ArcadeBench scores AI decisions against exact game-theoretic ground truth — no judge model, no rubric — and reports what fraction of a model’s metered capacity actually bought a correct move. Across 265 tournament matches, two thirds ended in a draw and a quarter of series had to be settled by coin toss. The result stopped telling us anything. The cost didn’t.
Tic-Tac-Toe is a draw under optimal play. As models stop making mistakes, match results converge on that draw and stop separating them. Two datasets, scored under identical method, show the collapse.
A bracket has to produce a winner, so a level series goes to a coin toss. Nineteen did — including a grand final, where a best-of-seven between GPT-5.6 Luna and Claude Sonnet 5 reached game seven still tied and the title went to heads. That is not a flaw in the tournament. It is what outcome-based ranking degrades into once players stop making mistakes.
Move regret doesn’t break the tie either: two players who both play perfectly both score zero, and that is the right answer. What separates them is what they spent. Over these same decisions the fraction of capacity that bought an optimal move ranged from 100.0% down to 57.3%.
Tic-Tac-Toe is small enough to search in full, so every legal move in every reachable position carries an exact value: +1 won, 0 drawn, -1 lost. The first quantity is how much value a decision destroyed.
Regret depends only on the position and the action chosen, so it is independent of the opponent’s strength — unlike a match result, which is substantially determined by the other player’s mistakes. The second quantity is the capacity settled for that decision, metered per request at the inference boundary and preserved as a receipt. Joining them gives the score.
Regret has a second reading. Because it is scored against exact truth, a non-zero value means the player’s implicit read of the position was not just worse but factually wrong — averaged across a player’s decisions, that makes move regret a rough, model-agnostic proxy for how often a player acted on a false belief about the position, the same failure mode ordinarily called hallucination in language-model output. It is a proxy, not a validated measurement: it is computed from the chosen action alone, not from any claim the player made.
Bounded to [0,1] for every match, so it compares directly across matches and models with no population-relative normalization.
One recorded AI-versus-AI match, 7 decisions, recomputed from the raw export. Every empty square is tinted by its exact value for the player to move. Step forward and watch where the capacity went.
Nothing above is read from a precomputed score. The positions come from the export’s state snapshots, the capacity from each request’s gateway receipt, and the values from an exact search of the game tree — the same code path you can run on your own export.
7 tournaments, 2,156 scored decisions across 2,167 recorded. Not a model ranking — per-model samples run 28 to 352 decisions on a solved game, and a bracket gives stronger players more games. See every model, with N on every row, or browse all 2,167 decisions.
ArcadeBench is the instrument. BarKade is the environment that produces the evidence, and its tournaments are why there is new evidence next week. Watch models play, then inspect any decision in the match.