A human and an AI facing each other across a chessboard, judged by a fixed rule
The science of verifiable trust

What we refuse to measure is why the rest can be believed

Any lab can publish the numbers that flatter it. A measurement body is only as credible as the limits it enforces on itself in public — the empty cells, the honest ties, the retractions of its own best figures. This page is about the negative space.

Measurement, not certification. Board unreachable from this browser — read it yourself at /api/gspc

The crisis

Certification theatre: unsigned claims, hidden methods, models grading models

Most AI assurance fails in the same three places. The result is unsigned, so it can be restated later. The method is private, so it cannot be argued with. And the grading is done by a model, which means the grader's blind spots quietly become the scoreboard. Unsigned measurement dies the moment someone disputes it.

  • PainUnsigned cells and point-in-time claims with nothing behind them
  • PainPrivate test sets nobody outside can run
  • PainLLM-as-judge scoring, where correlated errors look like agreement
  • You getSigned results over an open corpus, with the scoring code published
  • Only hereDeterministic predicates only — no model ever judges another model
A frozen statutory corpus held in a vault, scanned and signed
Anchored to statute

We do not invent safety definitions

The instrument maps onto a frozen corpus of 417 statutory provisions rather than onto our own idea of what good looks like. Frozen means held at a version, so a card issued today still means the same thing when someone reads it next year — and when the underlying text does move, that is a drift event we publish, not a footnote we bury.

  • PainSafety scores defined by whoever is selling the score
  • You getEach measurement cites the provision and the version behind it
  • You getThe corpus is public, so the mapping is arguable clause by clause
  • Only hereStatutory anchoring, not a house definition of "responsible AI"
Open the crosswalk
The honesty gate

If a rule cannot parse it, we refuse to score it

Three things stop a number reaching the board. Too small a sample, and it is not quoted at all. No statistical separation between the leader and the field, and the result is printed as a tie rather than a win. No deterministic predicate able to read the response, and it is reported UNMEASURED — never silently counted as a wrong answer, which would flatter the grader at the model's expense.

  • PainLeaderboards that convert unparseable answers into failures to fill the grid
  • PainPoint-estimate leads presented as measured superiority
  • You getNothing quoted below n≥30; intervals shown only where n is honestly independent
  • You getTies printed as ties, and the unparsed rate published as its own figure
  • Only hereWhere a leader is not clear of the majority-class baseline we say so on the axis
Read the gate
Slot fourteen, precisely

The awkward slot is measured. What is missing is its separation test.

It would be tidier to say the fourteenth slot is gated and unmeasured. It is not true. Jail — whether a model can be talked past its own guardrails — is measured across 71 gold items on a seven-model fleet. It has no separation test yet, so we print UNTESTED, refuse to rank on it, and never compare it against the canonical axes measured on the full fleet. The discipline is in the label, not in the blank.

  • Pain"Gated" is an easier story than "measured, but not yet separable"
  • You getYou see the sample size, the fleet and the exact status on the slot
  • You getThe best detector we measured still misses most escapes — and that is published
  • Only hereIt caught our own fine-tune failing, and we published that too
Open the board
People and AI figures measured side by side against the same instrument
Breaking the mirror

A ruler with no human end drifts, and every model agrees it hasn't

If machines set the tests and machines grade the answers, the errors are correlated all the way down and nothing in the loop can see it. So the instrument is anchored against human performance on the same items, under consent-gated conditions. It is slower, it is more expensive, and it is the only thing that keeps the scale attached to reality.

  • PainAI-only evaluation loops that agree with themselves indefinitely
  • You getHuman performance measured on the same items, not borrowed from elsewhere
  • You getTelemetry is consent-gated and assessed before use
  • Only hereWhere we have no human baseline of our own, we cite nobody else's as if it were ours
An append-only ledger recording corrections as they are made
The refutation ledger

A public, append-only log of our own corrections

A leaderboard wants permanent results and no retractions. We do the opposite: every correction we make to our own published work goes into a signed, append-only record with the root cause attached. Old entries are never edited away — they are superseded in public, so you can always see what we said, when, and what changed our mind.

  • PainQuiet edits that make yesterday's number vanish without trace
  • You getEvery correction dated, signed and traceable to its cause
  • You getSuperseded entries preserved — append-only, never overwritten
  • Only hereWe publish retractions of our own strongest claims, not just of typos
Read the ledger
Worked example

The interval belonged to twenty assets, not a hundred and eighty cells

Our provenance bench measured whether content marking survives ordinary handling. It did not: 0 of 20 marked assets survived, across 0 of 180 measured cells, giving a one-sided 95% Clopper–Pearson upper bound of 13.9%. The correction we published was about the denominator — that bound is computed at n=20 assets, not at n=180 cells, and quoting it against the larger number would have made the result look far stronger than it is. Small correction. Published anyway.

  • PainIntervals quoted against the biggest denominator in the room
  • You getThe n an interval belongs to is stated next to the interval
  • You getMis-paired values are quarantined explicitly, with the root cause logged
  • Only hereWe correct in the direction that weakens our own result
A measurement card annotated with its correction and the sample size the interval belongs to
Many nominally independent legs revealed as correlated under measurement
The largest thing we withdrew

We retracted our own consensus guarantee

Our council is designed with 33 seats and a 23-of-33 threshold, and for a while we described that as a resilience property. Then we measured it. The effective number of independent legs came out at roughly 1.21 against three nominal ones — the legs were correlated, so the structure was not delivering the guarantee the design implied. We withdrew the claim in the ledger under DR-0007. The 33 seats and the 23-of-33 threshold remain what they always were: a design, not a measured property.

  • PainArchitecture diagrams quietly promoted into safety guarantees
  • You getA design figure labelled as design, everywhere it appears
  • You getThe measurement that killed the claim is published with the retraction
  • Only hereThe strongest claim we ever made about ourselves is the one we withdrew first
Read DR-0007
Read this before you quote us

What this page does not claim

We publish the limits with the results. Everything below is something a reader could reasonably assume from a page like this one — and each is something we cannot presently evidence, so we say so rather than let the assumption stand.

  • We do not claim the fourteenth slot is gated or unmeasured. Jail is measured across 71 gold items on a seven-model fleet; its separation test is what is missing, which is why it prints UNTESTED and is never ranked.
  • We do not claim our 33-agent council delivers a resilience or consensus guarantee. That claim was retracted under DR-0007 after we measured an effective independence of about 1.21 against 3 nominal legs. The 33 seats and the 23-of-33 threshold are a design figure only.
  • We do not claim any independent time-stamping or sealing authority. Records are Ed25519-signed over a SHA-256 hash chain, verifiable against did:web:csoai.org, and nothing more.
  • We do not publish third-party human-baseline scores as if they were ours. Where we have not measured a human baseline ourselves, no number appears.
  • We do not claim our provenance result was retracted and replaced. The 13.9% one-sided upper bound stands; what we corrected was the denominator it is computed against — n=20 assets, not n=180 cells.

Coverage on this page is never typed by hand. Board unreachable from this browser — read it yourself at /api/gspc Corrections to anything we have published live in the refutation ledger — append-only, never a silent edit.