ArcadeBench
ModelNOptimalAvg regretJ / decisionEfficiency
gemini-3.1-pro-preview62100.0%0.00012,238
1.0000
claude-fable-549100.0%0.00048,156
1.0000
claude-opus-539100.0%0.00021,202
1.0000
gemini-3.7-flash238100.0%0.0001,433
1.0000
gemma-4-31b-it83100.0%0.0001,114
1.0000
grok-4.624699.6%0.00410,491
0.9864
gpt-5.6-sol13298.5%0.01510,960
0.9738
qwen3.8-27b7097.1%0.0296,864
0.9658
claude-sonnet-535297.4%0.0268,932
0.9640
gpt-5.6-terra13394.0%0.0606,526
0.9320
gpt-oss-120b15090.0%0.100377
0.9174
glm-5.34497.7%0.0233,455
0.8946
muse-spark-1.26692.4%0.0767,477
0.8903
gpt-5.6-luna28090.0%0.104943
0.8245
gemini-3.5-flash-lite10181.2%0.188899
0.7867
claude-haiku-4.58373.5%0.3013,307
0.6534
inkling2860.7%0.4642,129
0.5729

Efficiency Rating is the fraction of a model’s metered capacity that bought an optimal move, bounded to [0,1] per match. It says nothing about how often a model won — browse every scored decision for that, or rescore a match to see one worked end to end.