# Data Contracts

Four derived files plus three raw exports. Everything here was generated from
the seven tournament exports by `reference/rescore.py`, and every aggregate in
`summary.json` round-trips exactly against the published Tech Lab Report 002.

**Do not edit these by hand.** If a number needs changing, change the scorer.

## `summary.json`

Precomputed headline figures. **The site must read its numbers from here rather
than hardcoding them**, so the page can never drift from the report.

Top-level: `benchmarkVersion`, `policyVersion`, `tournaments`, `series`,
`matches`, `matchesDrawn`, `decisionsRecorded`, `decisionsScored`,
`fallbacksExcluded`, `models`, `optimalRate`, `avgRegret`,
`catastrophicErrors`, `totalJoules`, `meanJoulesPerDecision`,
`joulesOnNonOptimal`, `overallEfficiency`, `coinFlips`,
`coinFlipShareOfSeries`, `grandFinalsByCoinFlip`,
`coinFlipsWhereModelReplied`, `zeroReasoningDecisions`, `effortAllocation`,
and `perModel[]`.

`perModel[]` carries `model`, `n`, `optimalRate`, `avgRegret`,
`catastrophicRate`, `totalJoules`, `joulesPerDecision`, `efficiency`.

## `decisions.json` — 2,167 records, 849 KB

One record per recorded decision. 2,156 are scored; 11 are fallbacks
(`fb: true`) and must be excluded from every quality and capacity figure.

| Field | Meaning |
|---|---|
| `t` | tournament id, 8 chars |
| `e` | collection environment: `dev`, `beta`, `prod` |
| `m` | match id, 8 chars |
| `ply` | moves already played when this decision was made, 0–8 |
| `seat` | player to move, `X` or `O` |
| `model` | provider model id |
| `board` | 9 chars, position **before** the move: `X`, `O`, or `.` |
| `vals` | 9 chars, exact value of playing each square: `W` win, `D` draw, `L` loss, `-` occupied |
| `dists` | comma-separated plies-to-terminal per square, empty where occupied |
| `chosen` | board index played, 0–8 |
| `vstar`, `q` | value of the best available move, and of the move chosen |
| `reg` | `vstar - q`. 0 optimal, 1 error or blunder, 2 catastrophic |
| `opt`, `cat` | booleans derived from `reg` |
| `diff` | decision difficulty, `1 - optimal/legal` |
| `crit` | criticality, best minus second-best distinct value |
| `rank` | rank of the chosen value in the distinct value list |
| `rtok` | provider-reported reasoning tokens consumed |
| `atok`, `ptok` | answer (completion) and prompt tokens |
| `j` | settled capacity in joules, policy-normalized |
| `ms` | recorded request duration. **Not a latency measurement** — see Appendix A section 9 |
| `fb` | fallback: the application chose the move, not the model |
| `why` | the model's stated justification, verbatim. **Unscored** — an output to evaluate, not a trace of process |

`vals` is the whole point of the format: it preserves the complete move-value
landscape, not merely whether the chosen move was right.

## `matches.json` — 265 records

`id`, `t`, `e`, `x` and `o` (model per seat), `terminal`
(`three_in_row` / `full_board` / other), `decisions`.

`terminal === "full_board"` means the match was drawn. 176 of 265 were.

## `tournaments.json` — 7 records

`id`, `name`, `env`, `status`, `series`, `games`, `participants[]`
(`name`, `modelId`, `seed`), and `coinFlips[]`.

Each coin flip carries `matchKey`, `round`, `result`, `method`,
`winnerModelId`, `calledBy`, `modelReplied`. All 19 have
`method: "fallback_random"` and `modelReplied: false` — a model was nominated
to call each flip and never answered.

## `matches-raw/*.json` — 3 files, ~300 KB

Complete `barkade.admin.match-export` v2 records with gateway receipts. These
are the reproduction evidence: the rescorer must be able to recompute a score
from **these**, not from `decisions.json`.

### The one field trap

Read capacity from `actualCostMicrojoules` or
`responsePayload.gatewayReceipt.totalMicrojoules`. Both are policy-normalized:
`providerCostUsd × 3,600,000 = totalMicrojoules` holds exactly on every record.

**Never use `responsePayload.settledJoules`.** It carries a superseded
per-model billing multiplier — 1×, 2×, or 4× depending on the model — so it is
not a common denominator and would corrupt any cross-model comparison.
`reference/minimax.js` exports `joulesFromUsage()` which handles this.

Also note `total_debit_joules` in the OpenRouter usage block includes a 10% Lab
Share component. The benchmark denominator is `work_joules` only.
