◎ Portal← Back to the board

Brains Methodology

Everything behind every number on portal.ai/brains: the formulas, the routes, the guards, and the honest reasons some things stay private. Deep enough to reproduce the math from the public JSON alone.

v13 · v14 · v15 · v16 · v17 · v18 · v19 · v19.2 · v19.3Task mixAnchorIQFloorIFEval-miniBusiness suiteCodingLocal DeepSeekVelocity & deliveryIQ/$ONEMax stakesEmbeddingsRoute fidelityRelease integrityWhat stays privateDisclosures

v13 (shipped 2026-09-01): business-first without a silent rerank

v13 keeps ONE v1 and the legacy IQ formula unchanged while its new evidence fills. It AST-fingerprints the floor judges, refuses changed-method evidence, treats truncation as not-measured in every scoring probe, validates the declared weighted business mix, keeps prior fingerprints out of IQ and ONE, refuses stale price carry, and recomputes MAX STAKES only after the final ONE rows exist.

The new evidence layer is additive: six seeded Portal business tasks, twelve hidden-test coding canaries, and an aggregate-only local DeepSeek row. Until v17 the weighted business aggregate occupied the 30% IQ lane after a complete comparable sweep and coding stayed context-only; v17 (below) moves both lanes to the hourly run and into IQ. The local row stays context-only, and any ONE v2 proposal waits for two complete windows.

v14 (2026-09-01) adds Kimi K3 Max as the seventh commercial rail card under the same exact-receipt contract with a pinned route. A new card moves the board-relative noise band, so the model-set boundary resets trailing windows; ONE and IQ/$ carry forward while they refill.

v15 (2026-09-01) moves the Fable card to Fable 5.1 (Anthropic's same-day release) on the gateway's pinned 5.1 route. A card identity change is an epoch, not a model-set boundary: the Fable card's trailing windows reset and its ONE and IQ/$ show honest blanks until the lane re-measures, while every other card's history carries. The IQ/$ baseline stays the Fable card = 1.00×; 5.1 keeps 5's headline list price.

v16 (2026-09-01) rewrites the three prose judges of the floor battery (forecast-body rewrite, HD-manual explanation, pressure-free option frame) so they test the rule the probe states instead of a house vocabulary. The previous keyword lists rewarded echoing the prompt’s own words and accidental substrings while rejecting natural paraphrase: on the same run, three of the strongest cards failed the same two probes with fully compliant answers. Action is now any non-waiting imperative or modal-plus-infinitive advice; the option-frame probe names every constraint in its prompt (no imperative, no «должен» (“must”), no urgency word, one sentence). Judges stay deterministic code with no LLM; they were calibrated on sixteen real answers per probe before shipping. A judge change rotates the floor fingerprint, so floor windows refill from the change while ONE and IQ/$ carry and the chart keeps prior-methodology points for continuity.

v17 (2026-09-03) makes the operator’s own work the hourly measurement. Every run puts the six seeded business task shapes (fresh seed per run, the same slate for every card) and a rotating third of the twelve Portal coding canaries through each card’s exact route, and IQ becomes business 60% + coding 40%, each the mean over the last three runs (coding: the latest score per canary id; both lanes required — a missing lane blanks IQ for that hour instead of re-weighting). IFEval-mini stays measured once per card release and is shown, not weighted; the ten-probe floor stays as the hourly health gate and the reliability source. Velocity moves to the median wall time of one real task (10 = 4 s, thinking included). Every card receives the same one-line context with each task (no workspace, no tools, deliverable only) because the rail delivers the frontier cards as agents and an agent that «checks the workspace first» was scoring 0 on tasks it writes well. Two judges were re-calibrated on real answers the same night: the coding judge tests the function the prompt asked for (a demo line after a correct def no longer zeroes it; defensive raise is ordinary Python), and the narrative facets test meaning as morphology families instead of a few stems. Judges stay deterministic code, no LLM. The methodology boundary resets trailing windows; ONE and IQ/$ carry while they refill.

v18 (2026-09-03, operator-direct: «today’s 10 is the new 5») doubles the work behind every hourly lane. Each lane gains a HARD tier that asks for roughly twice the work — business: reconciliation of sixteen sources with five planted contradictions and two agreeing restatements, a five-cohort LTV/CAC chain with an annual plan, a rebate and a decision rule, a twelve-task dependency plan with earliest starts, makespan, zero-slack set and approvals, a seven-paragraph structured forecast under word bands and two banned words, a constrained outbound letter with eight facts and six bans, and a sixteen-field extraction from a long brief with three late revisions, four decoys and two unit conversions; partial credit on the hard tier is squared; code: twelve functions modeled on portal-hosted semantics (wallet lots, trial gates, fingerprint windows, PT→UTC with the DST rule, error taxonomy, redispatch, the Wilson interval, dependency ordering, exact USD micros, idempotent replay, and two in JavaScript). Each lane = 0.5 × base + 0.5 × hard, so by construction a card that only aces the pre-v18 tier reads 5. What the first measured runs showed, said plainly: the frontier cards ace the hard tier as well (hard business 9.4–10, hard code 8.1–10; composite 9.6–9.9), so the visible effect is compression near the top rather than a new midpoint — well-specified deterministic tasks are what these models do best, even at twice the work. A ceiling that actually moves needs reliability sampling (pass^k: a task counts only when every one of k independent samples passes) or the real-repository lane (fail-to-pass tests minted from actual portal-hosted commits); both are named here as the next step, not claimed as shipped. Same evening (2026-09-03 21:05 PT, operator-direct) the board went to six cards: Portal M3 500K Fast retired, and the Grok card moved from Grok 4.5 Fast to Grok 4.6 xhigh (effort xhigh, Fast off, standard rate card) under a card epoch — its pre-cutover components do not carry and its windows refill. IQ now shows the business lane alone; Code is its own column; the composite 0.8 × IQ + 0.2 × Code is the quality ONE and the badge rank on (the operator’s 80/20). The Python judge now allows loops, try/except and lambdas because it runs in a limited child process, and it judges immutability by behaviour instead of banning methods; JavaScript runs in a node vm with a per-call timeout. IFEval-mini leaves the board face for the drawer; a dated public index (Artificial Analysis Intelligence Index v4.1.1, manual pins with attribution) appears as context. The methodology boundary resets trailing windows; ONE and IQ/$ carry while they refill.

v19 (2026-09-04, operator-direct: «мерять качество коммуникации … хотя бы раз в час … чтобы я мог понять, что какая-то модель деградировала» — “measure the quality of communication … at least hourly … so I can tell which model degraded”) changes what IQ is made of. The saturation v18 measured was real (22 of 36 deterministic tasks read 10 for every card), so IQ moves off seeded JSON shapes onto communication quality: every run each card writes the same two short deliverables drawn from a pool of fourteen real shapes — an investor follow-up, a partner terms reply, a cofounder boundary message, a reseller pricing reply, a recruiting first touch, a founder message to owners after an incident, a per-person team doc, a recommendation, a warm-circle letter after a silence, a companion bot’s first touch and night line, forecast day-lines, a rendering into the reader’s language, a complaint reply, an operator report — on fictional composite recipients (public-safe, stable), in Russian and English. Four layers in the operator’s order: numbers (wall time, tokens, list cost, retries, detours), precision and completeness (a deterministic floor: language, length band, imperatives aimed at the reader, negation-planting, permission grammar, narration and scaffold leaks, connected-prose rhythm, required cells, numbers not in the brief), meaning and feeling (twenty-three anchored criteria — seen not seen-through, heard, safe, on their side, a mistake costs nothing, their own instrument, adult regard, clarity by cold-reader restatement, the practical turn, facts, completeness, native language, brevity, authority without servility, one ask, never a project, the instance not the clinic, night without invoice, held not spent, meaning-and-emotion in translation, report head, no order in the tail). The meaning and feeling layers are judged by a disclosed panel: three seats from three model lineages, never the writer’s own lineage, one grader call each covering every criterion with a verbatim quote before each verdict (yes 10 · mostly 7 · no 3 · breach 0), median across seats, splits published rather than averaged. Score = 0.45 feeling + 0.25 meaning + 0.20 precision + 0.10 craft, then caps (an invented fact ≤ 5.9, a kill-class breach ≤ 4.9, an imperative at the reader ≤ 6.9, a rhythm fail ≤ 7.5, a leak ≤ 7.2, the wrong language ≤ 3.0); the deliverable’s length is disclosed to the judges and never rewarded. The v18 business lane keeps running as a displayed floor tier (weight 0); Code is unchanged; the composite stays 0.8 × IQ + 0.2 × Code. New on every card: a degradation state against the card’s own trailing seven days (median/MAD band; one run under = watch, two consecutive = DEGRADED) — the board’s purpose is to show which brain degraded and how, hour by hour, without anyone having to taste anything. The methodology boundary resets trailing windows; ONE and IQ/$ carry while they refill. Judge cost stays small by construction: short deliverables, one grader call per seat, the cheapest legal lineage trio first, and no judge spend on a text the floor already capped.

v19.2 (2026-09-05, operator-direct: «мозг, а не рельс … считай нормально … мне нужны настоящие данные о продуктах разных компаний» — “the brain, not our rail … count properly … I need real data on the vendors’ products”) separates the two things v17 had deliberately fused. From 2026-09-04 13:22 PT our own serving pool behind Opus, Fable, GPT-5.6 and Kimi ran out of paid capacity for 10.5 hours; every hourly call to those four brains died in our gateway’s admission queue, and the board read it as four brains at IQ 2.0, Velocity 1.0 and ONE 0 for a day, then as IQ “recovering” from 2.0 to 7.8 in three hours as the zeros left the window. None of that was about any model. Since v19.2 an error whose body is our rail’s own text — the gateway’s capacity 429, its queue deadline, a no-capacity 503, an nginx 502 from a worker down, a connection refused — is class lane: not a measurement. A lane hour never enters IQ, Code, Velocity, Delivery or the degradation band; the card keeps its last measured windows, says “our lane failed this hour”, and shows the fact under Delivery as our lane N/M (failed calls of the last 24 hours, with our median queue wait in the tooltip). A vendor that cannot serve — overloaded, upstream failure, its own rate limit — stays provider and still reads as erroring: that is the product status this board exists to report. Two more corrections landed the same day. Velocity is now the brain’s time, the wall time of a real task minus our gateway’s admission-queue wait (the X-Portal-Queue-Ms header the gateway returns on every answer), so a starved pool can no longer make four frontier brains the slowest on the board. And the displayed IQ is the mean of the last six measured comms runs (twelve real shapes, about six daytime hours) instead of three — two shapes over three runs had swung a card from 5.9 to 9.7 hour to hour; the degradation state still judges every single run, so a real fall is flagged within the hour while the headline number stays a number. The FACTS criterion of the comms panel was recalibrated on evidence from the first day: judges had marked a recommendation’s “working with him meant I stopped checking the pipeline myself” as an invented fact on three of six cards; the writer’s own first-person experience is the writer’s voice, and only a checkable claim absent from the brief — a number, date, price, quote, a third party’s state or promise — is a fabrication. The criterion change moves the comms fingerprint, so the IQ window refills from the deploy; ONE and IQ/$ carry while it does.

v19.3 (2026-09-07, operator-direct: «у тебя появился доступ в Астру — добавь её в Brains и сделай тест» — “you now have access to Astra — add it to Brains and run the test”) adds the first row that does not ride the Portal rail. GPT-6 Astra is reachable only through the operator’s own ChatGPT-plan Codex seats (no API base URL, no key, identities never leave his Mac), so the row is produced in two halves: on the operator Mac a probe builds the exact task slate the board will use for the coming run hour — the twelve business shapes, the eight coding canaries of that run’s rotation, the two comms shapes — runs each request as one codex exec turn on the freshest seat (reasoning high, read-only sandbox, an empty working directory, notes and history off) and ships text, wall time and tokens under the board’s canonical request hash; on the box the canary answers every Astra call from that file and the same judges score it — deterministic business and coding scoring, and the comms panel with OpenAI excluded as a lineage (Grok, Kimi and Opus judge Astra; GPT-5.6 never does). What differs is disclosed: Codex’s own system prompt and tools ride every Astra turn while the rail cards get only the one-line work context; the floor battery is unseeded per run and cannot be pre-generated, so the row has no floor; there is no route receipt and no price, so the row is ONE-unranked and carries no IQ/$; Velocity is the Codex wall time per task with no queue in front of it. An hour the Mac did not generate reads «not measured» — never a zero.

The work mix that drives v4

A privacy-safe census of 201 completed runs from August 16-30 produced the following evidence. Counts and known wall time are both shown because neither is a perfect proxy for business value.

Work classRun shareKnown wall share
Product, coding, and operations28.4%27.6%
System evolution12.4%18.6%
Forecast and oracle9.5%14.0%
Person-depth work13.4%12.7%
Nexus and retrieval6.0%15.6%
High-stakes analysis12.9%3.2%
Growth automation6.0%2.9%

One operator arc may create multiple rows, and missing wall time is not imputed. The census informs a versioned task mix; it never changes weights silently between releases.

The anchor

10/10 = Peak Fable 5, June 9–12 2026 — the operator's reference window, an operator-declared composite constant: protocol floor 10/10, IFEval-mini 100%, every business and coding task clean in BOTH tiers (since v18 the pre-v18 ceiling maps to 5), a 20-second floor-battery velocity reference and, since v17, a 4-second single-task velocity reference. These are declared reference values, not retro-measured June results. The anchor never drifts with the market. When any brain sustains the 10.0 ceiling, the public JSON raises anchor.rebase_due and the anchor is re-based upward by an explicit operator decision — never silently.

IQ — the quality composite

IQ        = comms lane  = 0.45 feeling + 0.25 meaning + 0.20 precision + 0.10 craft, capped   # 2 real shapes per run, mean of the last 3 runs
Code      = coding lane = 0.5 × base + 0.5 × hard        # latest score per task over the last 3 runs (8 of 24 per run)
composite = 0.80 × IQ + 0.20 × Code                       # what ONE and the badge rank on (v19, 2026-09-04)
shown, not weighted: business floor tier (v18 seeded shapes) · floor (10 probes, hourly health) · IFEval-mini (50 items, once per release) · public index (dated pin)

Two weighted lanes since v19 — communication quality (IQ) and the operator’s code work (Code). IQ is scored by a deterministic floor plus a disclosed three-lineage judge panel with quoted evidence (the writer’s lineage never judges its own text); Code and the floor tier stay deterministic code, no LLM judge. A lane missing fresh, fingerprint-matched rows blanks the composite for that hour instead of re-weighting, so a stale component can never silently pollute a live score; a task, criterion, judge-prompt or weight change rotates the comms fingerprint and its window refills. The v18 business tier, IFEval-mini, the protocol floor and the public index are published beside IQ but do not enter it. "IQ" is Portal’s internal operational composite for real communication work, not a standardized intelligence test.

Protocol floor — 10 live probes, hourly

Every hour (07:00–22:00 PT) each brain answers one prompt containing ten protocol probes in one pass. The probes test the operator's real working protocols: instruction fidelity under competing constraints, exact-format echoes, negation handling, multilingual expectation checks (including Russian), refusal-shape correctness, and arithmetic-with-format traps. One point per probe passed by exact deterministic checks; probe parameters are randomized per run so the target moves, and regression fixtures pin every judge against past false-negative classes. The probe texts themselves stay private — see what stays private.

IFEval-mini — instruction following

A fixed 50-item Portal subset of the public IFEval task family, scored with the standard deterministic IFEval checkers (prompt-level pass), refreshed as a full sweep on release days; since v17 shown without entering IQ, and since v18 shown in the row drawer only. Published per card: pass/n, the percentage, and a Wilson 95% confidence interval:

center = (p + z²/2n) / (1 + z²/n),  half = z·√(p(1−p)/n + z²/4n²) / (1 + z²/n),  z = 1.96

The subset is fixed (item ids hash-pinned in the release manifest) so cards stay comparable across runs; it is a subset, so figures are not comparable to full public IFEval scores.

Public index — apples-to-apples context

Since v18 each commercial card shows the Artificial Analysis Intelligence Index v4.1.1 (artificialanalysis.ai) — a public composite of nine third-party evaluations (GDPval-AA v2, τ³-Banking, Terminal-Bench 2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR) on a 0-100 scale — as dated, manually verified pins with the variant the index measured and a source link. It is context for comparing the cards against public opinion; it never enters IQ, Code, the composite or ONE, and the Portal house lines carry no pin. Attribution: Artificial Analysis.

Business suite — real deliverables

v13 expands three fixed paragraphs into six seeded task shapes: grounded forecast, grounded Step-2, Nexus evidence synthesis, high-stakes truth reconciliation, unit economics, and Midas routing. At release time each task runs three times per brain, for 18 generations per card; since v17 the same six shapes also run once per card every hourly run with a fresh seed, and since v18 they are the BASE tier of the business lane (0.5 of it; the hard tier of six two-effort shapes is the other 0.5) — since v19 that lane is a displayed floor tier (weight 0); IQ is the comms lane. Deterministic hard checks remain its score source; narrative facets are compliance canaries (tested as morphology families since v17), not a claim that keyword checks can measure literary quality.

Every release seed changes facts, amounts, requests, and surface wording while preserving the paired exam across cards. Output schemas use neutral placeholders rather than answer-bearing examples. In the truth task, absence of a record remains unknown rather than becoming a fabricated negative. The release battery is published as evidence only when all 18 rows, the methodology fingerprint, and the private evidence hash match the installed release; the hourly lane enters IQ only through fingerprint-matched rows of the current judges.

Coding: daily canary now, agentic workbench next

The prior three-item Coding Shot saturated. v13 implements twelve seeded hidden-test fixtures across state preservation, credit arithmetic, HTTP error semantics, bounded retries, lease safety, idempotent billing, structured routing, dependencies, approvals, CSS, and HTML. It is labeled Coding Canaries because a completion benchmark is not a repository agent. Since v17 the canaries are also the hourly coding lane (four per run, all twelve every three runs) and since v18 its BASE tier — the hard tier adds four of twelve portal-hosted functions per run; that lane is the Code column, 20% of the composite; the once-a-day twelve-task shot stays published as release evidence.

The decision-grade coding lane is Portal-bench-Live: fresh private tasks minted from recent portal-hosted changes, executed in isolated worktrees, judged by tests that fail before and pass after, with pass rate, confidence interval, wall time, retries, and cost per accepted task. Public SWE scores remain context only.

Local DeepSeek row

The local DeepSeek V4 Flash IQ3 row (Portal PRO96, an aggregate receipt shipped from the box) was retired on 2026-09-07 at the operator’s request: it never entered IQ, ONE or IQ/$, its IFEval was a 19-item prefix of the 50-item subset, and a once-a-week aggregate had nothing to compare with an hourly board. The board carries the brains the operator can buy through the Portal rail or run through his own seats, all measured hourly by the same pipeline.

The row is deliberately ONE-unranked in v13. API cost is near zero but hardware time, electricity, and opportunity cost are not zero, so an infinite IQ-per-dollar value would be fabricated. The local row joins lane comparisons first; a versioned cost basis is required before it can enter ONE.

The serving identity disclosed above (quant, context window, model bytes, CPU-MoE placement) is exactly what the public JSON carries per row. Deeper hardware trade-off analysis — memory-bandwidth limits, alternative quant ladders, accelerator options — is operator infrastructure engineering and stays private, like every other pre-decision note behind this board.

Velocity & delivery

velocity = clamp(1, 10, 10 × 4000 / trailing-24h-median-task-wall-ms)     # v17: one real business/coding task, 10 = 4 s
(floor battery until the first work run lands: 10 × 20000 / trailing-24h-median-latency-ms)

The clock includes the full delivery pipeline as configured — session setup, routing, thinking time — because that is what a user actually waits for. Delivery 24h = clean first-attempt completions over measured attempts; provider errors, timeouts, retries, refusals, and detected classifier reroutes all count against it. Our-side configuration failures are excluded (they measure us, not the model).

IQ/$ — quality per dollar

IQ/$ = (IQ / workload$) ÷ (IQ_fable / workload$_fable)     → Fable 5.1 1M Max = 1.00×

The workload is identical per card and, since v17, it is the hourly work run itself — since v18 the floor health check + twelve business tasks (six base + six hard) + eight coding tasks (four base + four hard) of that run (trailing-24h median cost); IQ/$ uses the business lane (IQ) over that full two-tier run cost; IFEval-mini and the release battery are display-only evidence and no longer enter the basis. Priced at dated public list rates for each route (fast tiers where fast is used; the pricing table id is stamped in the JSON). Only the indexed ratio is published; absolute dollars stay private. A failed run is cheap because it failed — so the ratio exists only on clean runs, and a cheap failure can never top the column. Not a financial return.

ONE — the default-choice number

ONE answers the operator's actual question: "which brain do I run by default, right now?" It reproduces rational satisficing choice: quality differences too small to distinguish from the board's own measurement noise should cost nothing, while a real quality deficit must dominate any price advantage. Formula one_v1_2026-07-10, pre-registered, computed hourly from published components only — anyone can recompute it from brains.json:

Q_i    = mean IQ over the last 3 canary runs (clean runs only)
B      = clamp( median_i( (max−min of the same 3-run window) / 2 ),  0.5, 1.0 )
d_i    = max(0, max(Q) − Q_i − B)          # deficit BEYOND the live noise band
G_i    = exp( −(4·d_i / B)² )              # smooth quality gate, no cliff
econ_i = 1 + ln(1 + IQ/$_i)                # log: price advantage is bounded
spd_i  = max(0.3, velocity_i / 10)         # linear: users feel wall-time 1:1
del_i  = delivery% / 100                   # (1.0 when unmeasured)
ONE_i  = 100 × (G·econ·spd·del)_i / leader # leader = 100

Why each functional form

Constants (window 3, band clamp [0.5, 1.0], gate steepness 4, speed floor 0.3) are frozen in version one_v1; any change lands as a new version id with a changelog entry — never a silent retune. The release validator recomputes every published ONE from the same public components and fails the release on any mismatch.

★ Max stakes — the second lane

ONE is normalized over the publicly ranked cards: the Portal M2 house card is measured every hour but sits off the ranking by design and carries no ONE or badge (since 2026-09-03; before that the hidden card could hold the 100). One ranking cannot honestly serve two jobs. ONE ranks the default lane (volume work, customer deliverables). The gold ★ MAX STAKES badge marks the separate answer to a different question: "which brain when quality outranks cost and time?" It appears only when the highest recent dependable quality clears the second row by more than the live noise band; inside-noise ties receive no badge. It is a pointer, not a rank; it never moves ONE.

Embeddings — Portal vs a commercial vendor

Paired comparison on identical Portal-authored private fixtures: same documents, same relevance judgments, same 1024 dimensions, cosine similarity, same top-k, each engine at its vendor-recommended query/document settings. Four lanes: Code, English, Russian, Ukrainian.

Route fidelity — measuring the thing we claim to measure

Release integrity

What stays private — and the honest why