ArcadeBench
The evidence, in full

Explore the evidence.

Rescore a real match in your browser, browse every scored decision, or download everything behind the report — the scorer, the raw exports, and the field reference.

The rescorer

A score you cannot recompute is an assertion.

Everything below is derived in your browser from the raw evidence — state snapshots, gateway receipts, and the moves as played. Nothing is read from the scored dataset we published. That is the point: if this page read our own numbers back to you, it would prove nothing.

What is reconstructed, and from where

The reconstruction

A match export records the position after every move, the move played, and one usage record per AI request. Joining them recovers the decision as it was faced.

QuantityRead from
Position before a movesnapshots[stateVersion − 1].stateJson.board
Player to movethe same snapshot’s nextPlayer
Move playedmoves[].moveJson.index
Capacity settledaiMoveUsage[], joined on idempotencyKey
Reasoning tokensopenRouterRequests[], joined on idempotencyKey
Stated justificationresponsePayload.reasoning, verbatim and unscored
Value of every legal moveexact search of the game tree, in this browser
One field trap, stated plainly
Capacity is read from actualCostMicrojoules, or the gateway receipt’s totalMicrojoules where the older schema has no such field. Both are normalized to the stated policy baseline. Never responsePayload.settledJoules — it carries a superseded per-model billing multiplier of 1×, 2×, or 4×, so it is not a common denominator and would corrupt any comparison across models. The gateway usage block’s total_debit_joules is also wrong for this purpose: it includes a 10% Lab Share component, and the benchmark denominator is work_joules only.
Capacity, not energy
Joules here are converted from provider-reported usage cost under a stated policy baseline. They are a unit of purchased computational capacity, not measured electricity.
Pick your evidence

Recompute from a raw export

Load one of the three released match exports, or supply your own barkade.admin.match-export file — schema version 1 or 2. Your file is read in this browser tab and never uploaded; there is no server here to upload it to.

or paste an export
Solver self-test, run in this browser
4 of 4 passed. All nine openings draw; exactly four of eight replies to a centre opening hold the draw; a corner reply is one of them; an edge reply loses. If those failed, the solver would be wrong and so would every number on this site.
Checking the checker

If you don’t trust this page

The solver in your browser is a port of reference/rescore.py, and the two are verified to agree on all 2,156 scored decisions in the published set. The build refuses to ship if that agreement breaks, and it refuses to ship if four known Tic-Tac-Toe facts stop holding. Both programs are downloadable, so you can run the Python one against the same export and compare.

The evidence index

Every decision, and what each one was worth.

2,167 recorded decisions — 2,156 scored and 11 fallbacks, which are excluded from every published figure. Open any row for the full value landscape of the position, the capacity settled, and what the model said it was doing.

Find one decision

Filter, then cite

Opening a row puts its address in the URL, so a claim about one move can be checked by whoever reads it. The whole file loads once and is filtered in this tab.

Loading
Fetching decisions.json — 849 KB, fetched once, then filtered in this tab.
Two things deliberately absent
No latency comparison. The ms field is in the downloadable data for completeness, but recorded durations include waiting: the slowest requests produced fewer tokens than the median, and duration correlates only weakly with output volume (r = +0.15). It is not exposed as a sort or a filter because a speed comparison built on it would be wrong. Appendix A section 9 explains it.
No score on the stated justification. The why field is shown verbatim and left unscored. ArcadeBench measures decision quality; the reasoning process is not observed, and grading the text would claim otherwise.
Per model

What the instrument reports, by model label

This is not a ranking. Per-model samples run from 28 to 352 decisions on a solved game, and a bracket gives stronger players more games, so a model that survived longer contributed more decisions. N is on every row; the rows are ordered by N rather than by any score, so the table cannot be read as a league position.

ModelNOptimalAvg regretCatastrophicTotal capacityJ / decisionEfficiency
claude-sonnet-535297.4%0.026—3.14 MJ8,9320.9640
gpt-5.6-luna28090.0%0.1040.36%263.99 kJ9430.8245
grok-4.624699.6%0.004—2.58 MJ10,4910.9864
gemini-3.7-flash238100.0%0.000—340.96 kJ1,4331.0000
gpt-oss-120b15090.0%0.100—56.59 kJ3770.9174
gpt-5.6-terra13394.0%0.060—868.00 kJ6,5260.9320
gpt-5.6-sol13298.5%0.015—1.45 MJ10,9600.9738
gemini-3.5-flash-lite10181.2%0.188—90.85 kJ8990.7867
gemma-4-31b-it83100.0%0.000—92.44 kJ1,1141.0000
claude-haiku-4.58373.5%0.3013.61%274.50 kJ3,3070.6534
qwen3.8-27b7097.1%0.029—480.45 kJ6,8640.9658
muse-spark-1.26692.4%0.076—493.51 kJ7,4770.8903
gemini-3.1-pro-preview62100.0%0.000—758.74 kJ12,2381.0000
claude-fable-549100.0%0.000—2.36 MJ48,1561.0000
glm-5.34497.7%0.023—152.04 kJ3,4550.8946
claude-opus-539100.0%0.000—826.87 kJ21,2021.0000
inkling2860.7%0.4647.14%59.61 kJ2,1290.5729

The Efficiency Rating is a fraction of settled capacity, so it says nothing about how much was spent. Five models answered every decision optimally, so their ratings are identical at 1.0000 — and their capacity consumption differs by a factor of 43, from 1,114 J per decision to 48,156 J. That is exactly what this score does not measure: it reports the fraction of spend that was wasted, and none of them wasted any. A cost-premium layer is the obvious next scoring version, and this table is the argument for it.

Token count doesn’t track that spread, and sometimes runs against it. claude-fable-5 used fewer tokens per decision than gemma-4-31b-it — 847 against 1,277 — yet cost 43 times more, roughly a 65-fold difference in price per token. Across all 2,157 decisions carrying token counts, token count correlates with settled capacity at only r = 0.34: capacity is priced, not counted, so a token-count denominator can rank models backwards from a cost-based one.

Overall: 94.3% optimal, average regret 0.060, 6 catastrophic errors (0.28%), 14.29 MJ settled across 2,156 scored decisions, overall Efficiency Rating 0.9634.

On reasoning effort

Effort did not track difficulty

Over 2,155 decisions, reasoning effort shows no positive relationship with exact decision difficulty: r = -0.044 against difficulty and -0.026 against answer tokens, with a median of 46 reasoning tokens and 601 decisions that spent none at all. What effort does track, weakly, is how far into the game the position is (r = 0.230).

Read that as a correlation measured at one reasoning setting, not as a claim about what models can or cannot do. It is reported because the difficulty figure it is measured against is exact, which is unusual; it is not a scored dimension of the benchmark.

Downloads

All of it, in the form the figures came from.

17 files, 28.4 MB. The seven tournament exports are the primary evidence; the three match exports are the worked reproduction cases; the derived datasets are what the scorer produced from them. Nothing here has been trimmed or reformatted for publication.

What this is, and is not
A point-in-time snapshot, not a live feed. Benchmark version arcadebench-v0, capacity policy joule-capacity-v1, report version 1.2, generated from 7 tournament exports dated 2026-08-20. There is no API and no leaderboard endpoint behind this page; when new tournaments are run, a new snapshot is published rather than these files changing under you.
Capacity, not energy
The joule figures throughout are converted from provider-reported usage cost under a stated policy baseline. They measure purchased computational capacity. They are not measured electricity, and they do not measure intelligence, effort, or the value of a result.
Primary evidence

The seven tournament exports

Every figure in the report is drawn from these: 265 matches, 2,156 scored decisions, 14.29 MJ of settled capacity, and the 19 coin tosses. They are point-in-time admin exports from the BarKade database, 27.3 MB in total.

TournamentEnvSeriesGamesParticipantsCoin tosses
v2026.08.19-0x0000beta156087
v2026.08.19-0x0001beta73041
v2026.08.19-0x0001dev154362
v2026.08.19-0x0002dev154552
v2026.08.20-0x0000dev72433
v2026.08.20-0x0000prod73543
v2026.08.20-0x0001prod72831

All 19 tosses were resolved by fallback_random. A model was nominated to call each one and answered in 0 of 19 cases — that is, never, so every title that came down to a toss was settled by the fallback. The tournament exports record which model was asked each time.

Reproduction cases

Three complete match exports

These carry the full gateway receipt for every request, which is what makes a match independently rescorable. The rescorer reads these — not the derived scores — so a visitor can watch the numbers be produced rather than repeated.

The scorer

Two implementations that must agree

Run the Python one against any export above and compare it with what this site computes in your browser. If they disagree, one of them is wrong and the published figures are in question — which is the point of shipping both.

Derived datasets

What the scorer produced

Regenerable from the exports above. Do not edit them by hand — if a number needs changing, the scorer changes.

Field reference

Data contracts

data/schema.md, rendered in full. It is the authority on what every field means, including the two fields that look usable and are not.

Four derived files plus three raw exports. Everything here was generated from the seven tournament exports by reference/rescore.py, and every aggregate in summary.json round-trips exactly against the published Tech Lab Report 002.

Do not edit these by hand. If a number needs changing, change the scorer.

summary.json

Precomputed headline figures. The site must read its numbers from here rather than hardcoding them, so the page can never drift from the report.

Top-level: benchmarkVersion, policyVersion, tournaments, series, matches, matchesDrawn, decisionsRecorded, decisionsScored, fallbacksExcluded, models, optimalRate, avgRegret, catastrophicErrors, totalJoules, meanJoulesPerDecision, joulesOnNonOptimal, overallEfficiency, coinFlips, coinFlipShareOfSeries, grandFinalsByCoinFlip, coinFlipsWhereModelReplied, zeroReasoningDecisions, effortAllocation, and perModel[].

perModel[] carries model, n, optimalRate, avgRegret, catastrophicRate, totalJoules, joulesPerDecision, efficiency.

decisions.json — 2,167 records, 849 KB

One record per recorded decision. 2,156 are scored; 11 are fallbacks (fb: true) and must be excluded from every quality and capacity figure.

FieldMeaning
ttournament id, 8 chars
ecollection environment: dev, beta, prod
mmatch id, 8 chars
plymoves already played when this decision was made, 0–8
seatplayer to move, X or O
modelprovider model id
board9 chars, position before the move: X, O, or .
vals9 chars, exact value of playing each square: W win, D draw, L loss, - occupied
distscomma-separated plies-to-terminal per square, empty where occupied
chosenboard index played, 0–8
vstar, qvalue of the best available move, and of the move chosen
regvstar - q. 0 optimal, 1 error or blunder, 2 catastrophic
opt, catbooleans derived from reg
diffdecision difficulty, 1 - optimal/legal
critcriticality, best minus second-best distinct value
rankrank of the chosen value in the distinct value list
rtokprovider-reported reasoning tokens consumed
atok, ptokanswer (completion) and prompt tokens
jsettled capacity in joules, policy-normalized
msrecorded request duration. Not a latency measurement — see Appendix A section 9
fbfallback: the application chose the move, not the model
whythe model's stated justification, verbatim. Unscored — an output to evaluate, not a trace of process

vals is the whole point of the format: it preserves the complete move-value landscape, not merely whether the chosen move was right.

matches.json — 265 records

id, t, e, x and o (model per seat), terminal (three_in_row / full_board / other), decisions.

terminal === "full_board" means the match was drawn. 176 of 265 were.

tournaments.json — 7 records

id, name, env, status, series, games, participants[] (name, modelId, seed), and coinFlips[].

Each coin flip carries matchKey, round, result, method, winnerModelId, calledBy, modelReplied. All 19 have method: "fallback_random" and modelReplied: false — a model was nominated to call each flip and never answered.

matches-raw/*.json — 3 files, ~300 KB

Complete barkade.admin.match-export v2 records with gateway receipts. These are the reproduction evidence: the rescorer must be able to recompute a score from these, not from decisions.json.

The one field trap

Read capacity from actualCostMicrojoules or responsePayload.gatewayReceipt.totalMicrojoules. Both are policy-normalized: providerCostUsd × 3,600,000 = totalMicrojoules holds exactly on every record.

Never use responsePayload.settledJoules. It carries a superseded per-model billing multiplier — 1×, 2×, or 4× depending on the model — so it is not a common denominator and would corrupt any cross-model comparison. reference/minimax.js exports joulesFromUsage() which handles this.

Also note total_debit_joules in the OpenRouter usage block includes a 10% Lab Share component. The benchmark denominator is work_joules only.

The record itself

Method, results, and limits

The canonical account of what was measured and what it does and does not establish. This site never out-claims it.