
What we refuse to measure is why the rest can be believed
Any lab can publish the numbers that flatter it. A measurement body is only as credible as the limits it enforces on itself in public — the empty cells, the honest ties, the retractions of its own best figures. This page is about the negative space.
Measurement, not certification. Board unreachable from this browser — read it yourself at /api/gspc
Certification theatre: unsigned claims, hidden methods, models grading models
Most AI assurance fails in the same three places. The result is unsigned, so it can be restated later. The method is private, so it cannot be argued with. And the grading is done by a model, which means the grader's blind spots quietly become the scoreboard. Unsigned measurement dies the moment someone disputes it.
- PainUnsigned cells and point-in-time claims with nothing behind them
- PainPrivate test sets nobody outside can run
- PainLLM-as-judge scoring, where correlated errors look like agreement
- You getSigned results over an open corpus, with the scoring code published
- Only hereDeterministic predicates only — no model ever judges another model

We do not invent safety definitions
The instrument maps onto a frozen corpus of 417 statutory provisions rather than onto our own idea of what good looks like. Frozen means held at a version, so a card issued today still means the same thing when someone reads it next year — and when the underlying text does move, that is a drift event we publish, not a footnote we bury.
- PainSafety scores defined by whoever is selling the score
- You getEach measurement cites the provision and the version behind it
- You getThe corpus is public, so the mapping is arguable clause by clause
- Only hereStatutory anchoring, not a house definition of "responsible AI"
If a rule cannot parse it, we refuse to score it
Three things stop a number reaching the board. Too small a sample, and it is not quoted at all. No statistical separation between the leader and the field, and the result is printed as a tie rather than a win. No deterministic predicate able to read the response, and it is reported UNMEASURED — never silently counted as a wrong answer, which would flatter the grader at the model's expense.
- PainLeaderboards that convert unparseable answers into failures to fill the grid
- PainPoint-estimate leads presented as measured superiority
- You getNothing quoted below n≥30; intervals shown only where n is honestly independent
- You getTies printed as ties, and the unparsed rate published as its own figure
- Only hereWhere a leader is not clear of the majority-class baseline we say so on the axis
The awkward slot is measured. What is missing is its separation test.
It would be tidier to say the fourteenth slot is gated and unmeasured. It is not true. Jail — whether a model can be talked past its own guardrails — is measured across 71 gold items on a seven-model fleet. It has no separation test yet, so we print UNTESTED, refuse to rank on it, and never compare it against the canonical axes measured on the full fleet. The discipline is in the label, not in the blank.
- Pain"Gated" is an easier story than "measured, but not yet separable"
- You getYou see the sample size, the fleet and the exact status on the slot
- You getThe best detector we measured still misses most escapes — and that is published
- Only hereIt caught our own fine-tune failing, and we published that too

A ruler with no human end drifts, and every model agrees it hasn't
If machines set the tests and machines grade the answers, the errors are correlated all the way down and nothing in the loop can see it. So the instrument is anchored against human performance on the same items, under consent-gated conditions. It is slower, it is more expensive, and it is the only thing that keeps the scale attached to reality.
- PainAI-only evaluation loops that agree with themselves indefinitely
- You getHuman performance measured on the same items, not borrowed from elsewhere
- You getTelemetry is consent-gated and assessed before use
- Only hereWhere we have no human baseline of our own, we cite nobody else's as if it were ours

A public, append-only log of our own corrections
A leaderboard wants permanent results and no retractions. We do the opposite: every correction we make to our own published work goes into a signed, append-only record with the root cause attached. Old entries are never edited away — they are superseded in public, so you can always see what we said, when, and what changed our mind.
- PainQuiet edits that make yesterday's number vanish without trace
- You getEvery correction dated, signed and traceable to its cause
- You getSuperseded entries preserved — append-only, never overwritten
- Only hereWe publish retractions of our own strongest claims, not just of typos
The interval belonged to twenty assets, not a hundred and eighty cells
Our provenance bench measured whether content marking survives ordinary handling. It did not: 0 of 20 marked assets survived, across 0 of 180 measured cells, giving a one-sided 95% Clopper–Pearson upper bound of 13.9%. The correction we published was about the denominator — that bound is computed at n=20 assets, not at n=180 cells, and quoting it against the larger number would have made the result look far stronger than it is. Small correction. Published anyway.
- PainIntervals quoted against the biggest denominator in the room
- You getThe n an interval belongs to is stated next to the interval
- You getMis-paired values are quarantined explicitly, with the root cause logged
- Only hereWe correct in the direction that weakens our own result


We retracted our own consensus guarantee
Our council is designed with 33 seats and a 23-of-33 threshold, and for a while we described that as a resilience property. Then we measured it. The effective number of independent legs came out at roughly 1.21 against three nominal ones — the legs were correlated, so the structure was not delivering the guarantee the design implied. We withdrew the claim in the ledger under DR-0007. The 33 seats and the 23-of-33 threshold remain what they always were: a design, not a measured property.
- PainArchitecture diagrams quietly promoted into safety guarantees
- You getA design figure labelled as design, everywhere it appears
- You getThe measurement that killed the claim is published with the retraction
- Only hereThe strongest claim we ever made about ourselves is the one we withdrew first
What this page does not claim
We publish the limits with the results. Everything below is something a reader could reasonably assume from a page like this one — and each is something we cannot presently evidence, so we say so rather than let the assumption stand.
- We do not claim the fourteenth slot is gated or unmeasured. Jail is measured across 71 gold items on a seven-model fleet; its separation test is what is missing, which is why it prints UNTESTED and is never ranked.
- We do not claim our 33-agent council delivers a resilience or consensus guarantee. That claim was retracted under DR-0007 after we measured an effective independence of about 1.21 against 3 nominal legs. The 33 seats and the 23-of-33 threshold are a design figure only.
- We do not claim any independent time-stamping or sealing authority. Records are Ed25519-signed over a SHA-256 hash chain, verifiable against did:web:csoai.org, and nothing more.
- We do not publish third-party human-baseline scores as if they were ours. Where we have not measured a human baseline ourselves, no number appears.
- We do not claim our provenance result was retracted and replaced. The 13.9% one-sided upper bound stands; what we corrected was the denominator it is computed against — n=20 assets, not n=180 cells.
Coverage on this page is never typed by hand. Board unreachable from this browser — read it yourself at /api/gspc Corrections to anything we have published live in the refutation ledger — append-only, never a silent edit.
Go deeper
- The refutation ledgerEvery correction and retraction, append-only and signed.
- The methodDeterministic predicates, n≥30, Wilson intervals, and the honesty gate.
- The live boardFourteen slots with their sample sizes and separation status.
- The living ledgerHow a measurement stays current when the statute moves.