ONERank with ties (v21, the shape of Chatbot Arena, Arena-Hard and LiveBench): ONE = 100 × recent quality ÷ the leader’s, where quality = 0.8 IQ + 0.2 Code over the last 3 runs. A brain’s place counts only the brains whose quality interval lies entirely above its own (interval = 0.8 × the IQ window’s 95 % half-width); brains whose intervals overlap share a place, marked =. Among tied brains the cheaper (IQ/$) and faster one is listed first — economy and speed break ties, they never outvote quality. Nothing collapses to zero: the v11–v20 gate that printed ONE 0 for three frontier brains at IQ 6.6–7.0 is retired. The ★ MAX STAKES badge marks the separate max-stakes lane pick: highest recent dependable quality, for work where quality outranks cost and time.
The anchor10/10 = Peak Fable 5, Jun 9-12 2026 (the operator’s reference window), a declared composite ceiling. Since v19 (2026-09-04) the 10 means every communication criterion yes and a clean floor on the operator’s own shapes — the declared Peak Fable 5 experience of his communication work. When a brain sustains the ceiling, the anchor is re-based upward.
IQCommunication quality on the operator’s real shapes, measured every run (v22). Every brain works under the operator’s own house law, given to each card in the same system prompt: whole connected sentences, no em dashes, every checkable fact from the brief and nothing invented, no orders or planted negations aimed at the reader, the recipient’s specific situation named, one ask with a proposed date, every required cell present. Until v21 the board tested the bare brain from a two-sentence prompt, so the default habits that the operator’s rules normally remove (em dashes, staccato, imperatives) read as low IQ while the same brains under the same rules did his work well — the board now measures the brain as he uses it. Every run (twice a day) every brain writes the SAME four status prompts (an investor follow-up EN, a cofounder boundary message RU, a warm-circle letter after a silence RU, a companion bot’s first touch RU), on fictional composite recipients; the full run (once a day) adds two rotating breadth shapes from the other ten. First the deterministic floor (language, length, em dashes, orders at the reader, planted negations, leaks, rhythm, required cells, numbers not in the brief); then three judge seats from three model lineages, never the writer’s own, grade twenty-three anchored criteria with a verbatim quote before every verdict; splits are published. Score = 0.45 feeling + 0.25 meaning + 0.20 precision + 0.10 craft, then caps; length never rewarded. IQ = the fixed panel’s mean over the last six runs with its 95 % interval; STATUS compares each fixed prompt with the brain’s own trailing seven days on that prompt. A pairwise arena against a frozen house-brain reference (v21) remains available as an opt-in mode in the drawer methodology, not as the headline.
CodeThe operator’s code work, its own column: twelve base canaries + twelve hard functions modeled on portal-hosted semantics (Python and JavaScript, hidden tests in an isolated judge), eight per full run (once a day), latest score per task over its last three measured runs, 0.5 base + 0.5 hard. The composite 0.8 × IQ + 0.2 × Code is what ONE and the badge rank on. The v18 seeded business tier, IFEval-mini (once per release), the ten-probe floor (a health check every run) and the public index are shown, not weighted. ⚠ Saturated since v19: twenty of twenty-four canaries score ≥9.8 for every card; the column is a tie above 9.2 until a harder real-repository tier lands.
Velocity + deliverySpeed on the identical real tasks - the median wall time of one business or coding task as configured, thinking included, net of our own gateway's queue wait (10 = 4 s), not bare model inference. Delivery % counts the vendor's errors, retries and classifier reroutes over 24h; hours our own rail failed the call are shown as «our lane» and count against nobody's brain (v19.2, brain not rail).
IQ/$How much IQ a dollar buys on the identical measured workload (the full work run, once a day: two comms deliverables + twelve business tasks + eight coding tasks — the brain’s own generations, never the judge seats and, since v19.4, never the floor health probe), indexed to Fable 5.1 1M Max = 1.0×, at dated public list rates - not any provider's bill; the Astra row is priced at the vendor’s own API list for the same model with Codex’s harness tokens measured and excluded. Clean runs only, so a cheap failure can never top the column. Not a financial return.