
A benchmark you can memorise is not an instrument
Static tests leak into training data and then measure recall wearing reasoning's clothes. The way out is metrology: environments generated fresh for every evaluation, turn-based and seed-reproducible, scored by a fixed rule. This page sets out that doctrine — and marks clearly which parts of it run today and which are design.
Measurement, not certification. Board unreachable from this browser — read it yourself at /api/gspc
Static benchmarks leak, then flatter everyone
A fixed test published once ends up in the next training corpus, and from then on it measures memorisation. This is not a scandal, it is thermodynamics: any benchmark that stays still long enough becomes training data. Saturation and contamination are widely documented across the popular static suites — we do not restate other people's audit numbers here, but the direction of travel is not in dispute.
- PainYesterday's benchmark is tomorrow's training set
- PainA high score can mean recall rather than reasoning, and the score cannot tell you which
- You getAn instrument that generates a fresh instance every evaluation
- Only hereNothing fixed to memorise means nothing to contaminate

Procedural freshness is the whole argument
Procedurally generated interactive environments — novel per evaluation, with no stated goal and no instructions — force a system to work out the rules rather than recall them. That is the property worth building an instrument around, and it is the reason a game can be a measuring device rather than a leaderboard.
- PainLeaderboards reward whoever saw the test set most recently
- You getA fresh instance per evaluation, so scores are about reasoning under novelty
- You getNo instructions and no stated goal — the agent has to infer the rules
- Only hereContamination resistance is a design property, not a policy promise
The gap an interactive instrument can still see
The clearest existing example of this design is ARC-AGI-3, a set of novel, handcrafted, turn-based environments with no instructions. On the ARC Prize project's own published results, a human panel solves essentially all of them while frontier systems average well under one percent. We cite that as reported by the source, not measured here — we have not run it ourselves, and we do not put other people's numbers on our board.
- PainSaturated static suites can no longer separate strong systems from weak ones
- You getAn interactive, novel-per-eval instrument still has enormous discriminating power
- Only hereThird-party results are labelled reported and attributed — never absorbed into our own measurements

If it is not deterministic, it cannot be signed
This is the hard boundary between an entertaining arena and an instrument. Only turn-based, fixed-seed, tick-locked environments with an append-only action log can produce a result that someone else can reproduce — and only a reproducible result can be meaningfully signed. Real-time, stochastic or wall-clock-dependent environments are wonderful to watch and cannot be measurement.
- PainReal-time environments produce runs nobody can replay exactly
- You getFixed seed plus append-only action log means the run replays to the same result
- You getA deterministic scoring predicate, applied identically every time
- Only hereWe would rather leave an environment unmeasured than sign a result we cannot reproduce

Mechanical speed is not strategic superiority
The lesson from competitive game AI is that an unconstrained agent wins on execution and everyone reads it as reasoning. A comparative instrument therefore has to equalise the mechanics: capped action rates, matched observation — a locked camera rather than full-map injection — equal time controls on a tick-locked clock, and either a matched visual interface or a strictly typed API on both sides. Constrain the hands to measure the mind.
- PainSuperhuman input rates read as superhuman strategy
- PainFull-state access on one side quietly invalidates the comparison
- You getInformation parity and matched interfaces on both sides of every match
- Only hereCalibration constraints are published with the result, so the fairness is auditable
Self-play does not make a good partner
An agent trained entirely against copies of itself learns to expect an optimal mechanical partner, and then fails with real people who hesitate, improvise and change their minds. Cooperative environments under permissive licences — the Overcooked-AI line of work is the canonical one — make that failure measurable. Alignment measured only in AI-versus-AI conditions is measuring the wrong thing.
- PainSelf-play agents that collapse the moment a human joins the team
- You getCoordination measured with real human partners, not simulated ones
- Only hereHuman-in-the-loop is a condition of the measurement, not an optional extra


Whoever runs the game must not be whoever scores it
This is the part that makes the rest safe. MEOK operates and hosts the play, and may monetise it. Council of AI measures the same environments as frozen instruments, publishes the results, and is never paid by any party it ranks. The operator and the measurer are separate by construction, exactly as they are everywhere else on this site.
- PainAn arena operator that also sets the scores is a promoter, not an instrument
- You getThe environments are frozen and published, so the measurement is reproducible off-platform
- Only hereWe take no money from anything we rank — including anything we play against
The apparatus is doctrine. The board is live.
Being precise about this matters more than the pitch. What runs today is the fourteen-slot board, the signed cards, the public verification endpoint and the swarm axis — a protocol bank measuring multi-step agent behaviour, with its effective-n caveat published. What is design is the games apparatus itself: the proposed efficiency predicate, the rating engine, the calibrated arena. And auto-generating fresh environments from existing benchmarks is a live research direction, not a product we are promising — the moment a language model enters the scoring loop, our first design law is broken.
- PainRoadmap architecture presented as shipped capability
- You getA published doctrine you can hold us to as it gets built
- You getEnvironments named as candidates are permissively licensed, so anyone can rebuild the instrument
- Only hereDesign is labelled design on every surface, including this one
What this page does not claim
We publish the limits with the results. Everything below is something a reader could reasonably assume from a page like this one — and each is something we cannot presently evidence, so we say so rather than let the assumption stand.
- We do not claim the games apparatus is built or measuring anything today. It is a published doctrine. What runs is the fourteen-slot board, the signed cards, the verification endpoint and the swarm axis.
- We do not claim an automatic generator that turns existing benchmarks into fresh environments. That is a research direction, and any design that puts a language model into the scoring loop is ruled out by our first design law: no model judges another model.
- We do not claim a fifteen-axis board. The board is fourteen slots; slot 15 is measured in-lane only, is never board-quotable, and is never counted in any total.
- We do not claim the ARC-AGI-3 results as our own measurement. They are third-party figures from the ARC Prize project, reported here and attributed, and they are not on our board.
- We do not publish other projects' saturation, contamination or timing figures as if we had measured them. The arguments on this page stand without borrowed numbers.
- We do not claim to host or measure any environment whose licence forbids it. Only permissively licensed environments are named as candidates for a frozen instrument.
Coverage on this page is never typed by hand. Board unreachable from this browser — read it yourself at /api/gspc Corrections to anything we have published live in the refutation ledger — append-only, never a silent edit.