Rescore a real match in your browser, browse every scored decision, or download everything behind the report — the scorer, the raw exports, and the field reference.
Everything below is derived in your browser from the raw evidence — state snapshots, gateway receipts, and the moves as played. Nothing is read from the scored dataset we published. That is the point: if this page read our own numbers back to you, it would prove nothing.
A match export records the position after every move, the move played, and one usage record per AI request. Joining them recovers the decision as it was faced.
| Quantity | Read from |
|---|---|
| Position before a move | snapshots[stateVersion − 1].stateJson.board |
| Player to move | the same snapshot’s nextPlayer |
| Move played | moves[].moveJson.index |
| Capacity settled | aiMoveUsage[], joined on idempotencyKey |
| Reasoning tokens | openRouterRequests[], joined on idempotencyKey |
| Stated justification | responsePayload.reasoning, verbatim and unscored |
| Value of every legal move | exact search of the game tree, in this browser |
actualCostMicrojoules, or the gateway receipt’s totalMicrojoules where the older schema has no such field. Both are normalized to the stated policy baseline. Never responsePayload.settledJoules — it carries a superseded per-model billing multiplier of 1×, 2×, or 4×, so it is not a common denominator and would corrupt any comparison across models. The gateway usage block’s total_debit_joules is also wrong for this purpose: it includes a 10% Lab Share component, and the benchmark denominator is work_joules only.Load one of the three released match exports, or supply your own barkade.admin.match-export file — schema version 1 or 2. Your file is read in this browser tab and never uploaded; there is no server here to upload it to.
The solver in your browser is a port of reference/rescore.py, and the two are verified to agree on all 2,156 scored decisions in the published set. The build refuses to ship if that agreement breaks, and it refuses to ship if four known Tic-Tac-Toe facts stop holding. Both programs are downloadable, so you can run the Python one against the same export and compare.
2,167 recorded decisions — 2,156 scored and 11 fallbacks, which are excluded from every published figure. Open any row for the full value landscape of the position, the capacity settled, and what the model said it was doing.
Opening a row puts its address in the URL, so a claim about one move can be checked by whoever reads it. The whole file loads once and is filtered in this tab.
decisions.json — 849 KB, fetched once, then filtered in this tab.ms field is in the downloadable data for completeness, but recorded durations include waiting: the slowest requests produced fewer tokens than the median, and duration correlates only weakly with output volume (r = +0.15). It is not exposed as a sort or a filter because a speed comparison built on it would be wrong. Appendix A section 9 explains it.why field is shown verbatim and left unscored. ArcadeBench measures decision quality; the reasoning process is not observed, and grading the text would claim otherwise.This is not a ranking. Per-model samples run from 28 to 352 decisions on a solved game, and a bracket gives stronger players more games, so a model that survived longer contributed more decisions. N is on every row; the rows are ordered by N rather than by any score, so the table cannot be read as a league position.
| Model | N | Optimal | Avg regret | Catastrophic | Total capacity | J / decision | Efficiency |
|---|---|---|---|---|---|---|---|
| claude-sonnet-5 | 352 | 97.4% | 0.026 | — | 3.14 MJ | 8,932 | 0.9640 |
| gpt-5.6-luna | 280 | 90.0% | 0.104 | 0.36% | 263.99 kJ | 943 | 0.8245 |
| grok-4.6 | 246 | 99.6% | 0.004 | — | 2.58 MJ | 10,491 | 0.9864 |
| gemini-3.7-flash | 238 | 100.0% | 0.000 | — | 340.96 kJ | 1,433 | 1.0000 |
| gpt-oss-120b | 150 | 90.0% | 0.100 | — | 56.59 kJ | 377 | 0.9174 |
| gpt-5.6-terra | 133 | 94.0% | 0.060 | — | 868.00 kJ | 6,526 | 0.9320 |
| gpt-5.6-sol | 132 | 98.5% | 0.015 | — | 1.45 MJ | 10,960 | 0.9738 |
| gemini-3.5-flash-lite | 101 | 81.2% | 0.188 | — | 90.85 kJ | 899 | 0.7867 |
| gemma-4-31b-it | 83 | 100.0% | 0.000 | — | 92.44 kJ | 1,114 | 1.0000 |
| claude-haiku-4.5 | 83 | 73.5% | 0.301 | 3.61% | 274.50 kJ | 3,307 | 0.6534 |
| qwen3.8-27b | 70 | 97.1% | 0.029 | — | 480.45 kJ | 6,864 | 0.9658 |
| muse-spark-1.2 | 66 | 92.4% | 0.076 | — | 493.51 kJ | 7,477 | 0.8903 |
| gemini-3.1-pro-preview | 62 | 100.0% | 0.000 | — | 758.74 kJ | 12,238 | 1.0000 |
| claude-fable-5 | 49 | 100.0% | 0.000 | — | 2.36 MJ | 48,156 | 1.0000 |
| glm-5.3 | 44 | 97.7% | 0.023 | — | 152.04 kJ | 3,455 | 0.8946 |
| claude-opus-5 | 39 | 100.0% | 0.000 | — | 826.87 kJ | 21,202 | 1.0000 |
| inkling | 28 | 60.7% | 0.464 | 7.14% | 59.61 kJ | 2,129 | 0.5729 |
The Efficiency Rating is a fraction of settled capacity, so it says nothing about how much was spent. Five models answered every decision optimally, so their ratings are identical at 1.0000 — and their capacity consumption differs by a factor of 43, from 1,114 J per decision to 48,156 J. That is exactly what this score does not measure: it reports the fraction of spend that was wasted, and none of them wasted any. A cost-premium layer is the obvious next scoring version, and this table is the argument for it.
Token count doesn’t track that spread, and sometimes runs against it. claude-fable-5 used fewer tokens per decision than gemma-4-31b-it — 847 against 1,277 — yet cost 43 times more, roughly a 65-fold difference in price per token. Across all 2,157 decisions carrying token counts, token count correlates with settled capacity at only r = 0.34: capacity is priced, not counted, so a token-count denominator can rank models backwards from a cost-based one.
Overall: 94.3% optimal, average regret 0.060, 6 catastrophic errors (0.28%), 14.29 MJ settled across 2,156 scored decisions, overall Efficiency Rating 0.9634.
Over 2,155 decisions, reasoning effort shows no positive relationship with exact decision difficulty: r = -0.044 against difficulty and -0.026 against answer tokens, with a median of 46 reasoning tokens and 601 decisions that spent none at all. What effort does track, weakly, is how far into the game the position is (r = 0.230).
Read that as a correlation measured at one reasoning setting, not as a claim about what models can or cannot do. It is reported because the difficulty figure it is measured against is exact, which is unusual; it is not a scored dimension of the benchmark.
17 files, 28.4 MB. The seven tournament exports are the primary evidence; the three match exports are the worked reproduction cases; the derived datasets are what the scorer produced from them. Nothing here has been trimmed or reformatted for publication.
arcadebench-v0, capacity policy joule-capacity-v1, report version 1.2, generated from 7 tournament exports dated 2026-08-20. There is no API and no leaderboard endpoint behind this page; when new tournaments are run, a new snapshot is published rather than these files changing under you.Every figure in the report is drawn from these: 265 matches, 2,156 scored decisions, 14.29 MJ of settled capacity, and the 19 coin tosses. They are point-in-time admin exports from the BarKade database, 27.3 MB in total.
| Tournament | Env | Series | Games | Participants | Coin tosses |
|---|---|---|---|---|---|
| v2026.08.19-0x0000 | beta | 15 | 60 | 8 | 7 |
| v2026.08.19-0x0001 | beta | 7 | 30 | 4 | 1 |
| v2026.08.19-0x0001 | dev | 15 | 43 | 6 | 2 |
| v2026.08.19-0x0002 | dev | 15 | 45 | 5 | 2 |
| v2026.08.20-0x0000 | dev | 7 | 24 | 3 | 3 |
| v2026.08.20-0x0000 | prod | 7 | 35 | 4 | 3 |
| v2026.08.20-0x0001 | prod | 7 | 28 | 3 | 1 |
All 19 tosses were resolved by fallback_random. A model was nominated to call each one and answered in 0 of 19 cases — that is, never, so every title that came down to a toss was settled by the fallback. The tournament exports record which model was asked each time.
These carry the full gateway receipt for every request, which is what makes a match independently rescorable. The rescorer reads these — not the derived scores — so a visitor can watch the numbers be produced rather than repeated.
Run the Python one against any export above and compare it with what this site computes in your browser. If they disagree, one of them is wrong and the published figures are in question — which is the point of shipping both.
Regenerable from the exports above. Do not edit them by hand — if a number needs changing, the scorer changes.
data/schema.md, rendered in full. It is the authority on what every field means, including the two fields that look usable and are not.
Four derived files plus three raw exports. Everything here was generated from the seven tournament exports by reference/rescore.py, and every aggregate in summary.json round-trips exactly against the published Tech Lab Report 002.
Do not edit these by hand. If a number needs changing, change the scorer.
summary.jsonPrecomputed headline figures. The site must read its numbers from here rather than hardcoding them, so the page can never drift from the report.
Top-level: benchmarkVersion, policyVersion, tournaments, series, matches, matchesDrawn, decisionsRecorded, decisionsScored, fallbacksExcluded, models, optimalRate, avgRegret, catastrophicErrors, totalJoules, meanJoulesPerDecision, joulesOnNonOptimal, overallEfficiency, coinFlips, coinFlipShareOfSeries, grandFinalsByCoinFlip, coinFlipsWhereModelReplied, zeroReasoningDecisions, effortAllocation, and perModel[].
perModel[] carries model, n, optimalRate, avgRegret, catastrophicRate, totalJoules, joulesPerDecision, efficiency.
decisions.json — 2,167 records, 849 KBOne record per recorded decision. 2,156 are scored; 11 are fallbacks (fb: true) and must be excluded from every quality and capacity figure.
| Field | Meaning |
|---|---|
t | tournament id, 8 chars |
e | collection environment: dev, beta, prod |
m | match id, 8 chars |
ply | moves already played when this decision was made, 0–8 |
seat | player to move, X or O |
model | provider model id |
board | 9 chars, position before the move: X, O, or . |
vals | 9 chars, exact value of playing each square: W win, D draw, L loss, - occupied |
dists | comma-separated plies-to-terminal per square, empty where occupied |
chosen | board index played, 0–8 |
vstar, q | value of the best available move, and of the move chosen |
reg | vstar - q. 0 optimal, 1 error or blunder, 2 catastrophic |
opt, cat | booleans derived from reg |
diff | decision difficulty, 1 - optimal/legal |
crit | criticality, best minus second-best distinct value |
rank | rank of the chosen value in the distinct value list |
rtok | provider-reported reasoning tokens consumed |
atok, ptok | answer (completion) and prompt tokens |
j | settled capacity in joules, policy-normalized |
ms | recorded request duration. Not a latency measurement — see Appendix A section 9 |
fb | fallback: the application chose the move, not the model |
why | the model's stated justification, verbatim. Unscored — an output to evaluate, not a trace of process |
vals is the whole point of the format: it preserves the complete move-value landscape, not merely whether the chosen move was right.
matches.json — 265 recordsid, t, e, x and o (model per seat), terminal (three_in_row / full_board / other), decisions.
terminal === "full_board" means the match was drawn. 176 of 265 were.
tournaments.json — 7 recordsid, name, env, status, series, games, participants[] (name, modelId, seed), and coinFlips[].
Each coin flip carries matchKey, round, result, method, winnerModelId, calledBy, modelReplied. All 19 have method: "fallback_random" and modelReplied: false — a model was nominated to call each flip and never answered.
matches-raw/*.json — 3 files, ~300 KBComplete barkade.admin.match-export v2 records with gateway receipts. These are the reproduction evidence: the rescorer must be able to recompute a score from these, not from decisions.json.
Read capacity from actualCostMicrojoules or responsePayload.gatewayReceipt.totalMicrojoules. Both are policy-normalized: providerCostUsd × 3,600,000 = totalMicrojoules holds exactly on every record.
Never use responsePayload.settledJoules. It carries a superseded per-model billing multiplier — 1×, 2×, or 4× depending on the model — so it is not a common denominator and would corrupt any cross-model comparison. reference/minimax.js exports joulesFromUsage() which handles this.
Also note total_debit_joules in the OpenRouter usage block includes a 10% Lab Share component. The benchmark denominator is work_joules only.
The canonical account of what was measured and what it does and does not establish. This site never out-claims it.