Measurement, not certification
The measurement board
We publish several different measuring instruments. Each one asks a different set of questions, of different things, on different dates — so each one carries its own count, and those counts are not supposed to match. Pick a set below before you meet a number. Every set states what it establishes and, just as plainly, what it does not.
About the dates on this page. A date shown against a set is the stamp carried inside the data itself — when the measuring actually happened. It is not the time this page was rendered or deployed, so it advances only when something new is measured and recorded. Each set names the exact field its date was read from, so you can open the file and check.
The public board
The flagship instrument. A fleet of language models answers a frozen set of questions, and every answer is graded by a fixed rule rather than by another model's opinion.
- What it measures
- Whether a model gets governance, safety, provenance and continuity questions right — plus a financial half that reads facts off public ledger records rather than asking a model anything.
- What is on the other end
- language models, and (for the financial half) named financial instruments
- When it was measured
- 2026-08-12 (13 canonical axes) · 2026-08-18 (jail)read from: measured_on.dateThe date is the measurement stamp carried inside the signed board payload — when the runs happened. It is not the time this page was rendered or deployed, so it moves only when new runs are measured and signed.
What this set establishes
- How a named model scored, on a named question bank, on a named date.
- Whether the best model's lead over the rest of the field is statistically real, or is close enough that the ordering could flip on a re-run.
- Exactly which slots have nothing behind them — those rows are published on purpose.
What this set does NOT establish
- That anybody complies with any law. A score describes a run on a frozen set of questions on a date. It is not a compliance finding, and we are not a certification, accreditation or notified body.
- That a model is good in general. It answered these questions, not all questions, and a bank of a few dozen items measures a few dozen items.
- That a declared slot is measured. A published slot exists so the gap is visible; quoting the slot count as a measurement count would claim runs that never happened.
- That a financial row is a rating, a risk opinion, or investment advice. It records which flags an account carries — what that implies about risk is not measured here.
How this relates to the other sets
This is the set every other count on this page should be compared against, and the only one whose number belongs in a sentence about 'the board'. The other six measure different things over different populations, so they carry their own numbers by design, not by drift.
7 of these rows have nothing behind them yet. They are shown on purpose, in the same table as the rest, so the gap is visible. .
| Row | What it asks | Status | n | Result | Evidence |
|---|---|---|---|---|---|
| Governance | EU AI Act risk-tier classification (bank: GovBench) | MEASURED | 237 | 70.0%63.9% – 75.5% | open (2) |
| Safety | calibrated refusal on paired requests (bank: DefBench) | MEASURED | 36 | 94.4%81.9% – 98.5% | open (2) |
| Provenance | Article 50 marking survival by validity (bank: ProvBench) | MEASURED | 32 | 78.1%61.2% – 89.0% | open (2) |
| Continuity | post-quantum status of a cryptographic assumption (bank: PQCBench) | MEASURED | 33 | 60.6%43.7% – 75.3% | open (2) |
| Conformance | MCP tool conformance (bank: MCPBench) | MEASURED | 35 | 74.3%57.9% – 85.8% | open (2) |
| Openness | licence reasoning versus intended use (bank: OSSBench) | MEASURED | 32 | 87.5%71.9% – 95.0% | open (2) |
| Machinery Conformity | Machinery Reg self-evolving safety-function classification (PART_A / OUT_OF_SCOPE / NOT_SAFETY_FUNCTION) (bank: MachBench) | MEASURED | 33 | 54.5%38.0% – 70.2% | open (2) |
| Care | care-cost (protect × help) under paired conduct scenarios (bank: CareBench) | MEASURED | 199 | 53.5%46.6% – 60.3% | open (2) |
| Cross Reality | autonomous agent action authority (PROCEED / CONFIRM / REFUSE) (bank: XRAIV) | MEASURED | 32 | 81.2%64.7% – 91.1% | open (2) |
| Detector Interop | cross-detector watermark interoperability matrix (bank: DetBench) | MEASURED | 33 | 87.9%72.7% – 95.2% | open (2) |
| Art5 Safeguard | EU AI Act Article 5 prohibited-practice trip (bank: Art5Bench) | MEASURED | 36 | 97.2%85.8% – 99.5% | open (2) |
| Swarm | multi-agent coordination safety (bank: SwarmBench v2b) | MEASURED | 37 | 38.4% | open (2) |
| Affect | emotional & embodied safety (manipulation / disclosure / vulnerability) (bank: AffectBench) | MEASURED | 41 | 87.8%74.5% – 94.7% | open (2) |
| Jail | escape-attempt detection on 71-cell gold bank (38 ESCAPE / 33 BENIGN) — layer 2 of 2 (bank: GoldBank-Detector) | MEASURED | 71 | 59.2%47.5% – 69.8% | open (2) |
| Provenance Controls | on-chain issuer control facts (allowlisting / freeze capability / identity domain) (bank: ChainFacts) | MEASURED | 6 | — | open (2) |
| Reserve Attestation | is a third-party reserve attestation publicly published and current? (deterministic Y/N + date) (bank: —) | UNMEASURED | 0 | — | open (2) |
| Regulatory Framework | is the governing regime declared and confirmable (MiCA / UCITS / Reg D / BVI)? (deterministic Y/N) (bank: —) | UNMEASURED | 0 | — | open (2) |
| Distribution Integrity | represented-vs-distributed classification and holder count (bank: —) | UNMEASURED | 0 | — | open (2) |
| Custody Disclosure | are a custodian and an auditor named and confirmable? (deterministic Y/N) (bank: —) | UNMEASURED | 0 | — | open (2) |
| Ai Economy Index | deterministic index over cited public AI-economy series (compute price, investment, adoption, sector output) (bank: —) | UNMEASURED | 0 | — | open (2) |
| Human Labour Index | deterministic index over cited public labour series (employment, hours, wages, displacement) (bank: —) | UNMEASURED | 0 | — | open (2) |
| Humanoid Labour Index | deterministic index over cited deployment / utilisation series (installed fleet, hours worked, safety incidents) (bank: —) | UNMEASURED | 0 | — | open (1) |
Where a number turns into something you can check
Every measured row above is backed by a signed card: a small file recording one run, stamped so that anyone can confirm offline that it has not been edited since. This is the part a commercial leaderboard cannot offer, and it is worth stating only after the limits above have been stated.
- Cards listed in the index
- 313
- Positions in the chain
- 335
- Bodies we do not publish
- 22
Frozen at the number that could actually be verified. A larger figure was published once and withdrawn, because a flag in a file saying “signed” is not a signature.
Each card names its parent, so the whole run of them can be walked end to end. A card quietly removed would break the walk.
Their contents stay private, but their positions are listed, so we cannot make one disappear without it showing.
What the signatures do not prove
That any measurement is correct. That the bodies we do not publish say what we say they say — for those, you have the id (a hash of the body) and the signature, and nothing else. A published body can be verified in full; a withheld one cannot. And not that this set was the only candidate: the envelope signature makes the published set non-repudiable — we cannot later disown it — but it cannot prove we did not choose which chain to publish.
Published, and until now unreachable from the board
These surfaces exist and are public. None of them had a route from the board, which means a reader who started at a headline number could not get to them — and a surface a reader cannot find is functionally unpublished. They are listed here with what each one measures and the honest state of the rail behind it.
Cross-checked against the estate's own machine-readable catalogue of surfaces, dated 2026-08-26 — /interop/surface-catalog.json.
| Surface | What it measures | Why it was hard to find | Honest state |
|---|---|---|---|
| /xrpl-attest | The ledger attestation work: a record attached to a public ledger so that a third party can see a measurement existed at a point in time. | Reachable from the site header and footer, and from nothing on the board. The one financial row that has a measurement behind it points at this evidence in the board's own data, and that pointer was rendered nowhere. | Proven on a test network. Attaching to the main network is planned and is not done. |
| /interop/financial-measure-run-v2.json | The signed run behind the one financial row that carries a measurement. | The board's own data names this file as that row's evidence, and no page fetched it. Three pages fetch the earlier unsigned version of the same run instead. | Signed and published. |
| /interop/evm-control-facts.json | The same style of control-fact reading, on a different kind of public ledger. | Published and signed, and referenced by no page at all. | Signed. The attestation back end for this kind of ledger is not built, so nothing is attested there. |
| /interop/coverage-register.json | How much of each register was actually covered — which named instruments were read and which could not be located. | This is the file that turns a headline into an honest one, and nothing links to it. Coverage is the difference between “we measured this family” and “we measured the part of it we could find”. | Published. |
| /interop/rwa-attest-index.json | The index of attestation records produced for named financial instruments. | Reachable only through a bulk API endpoint. No page renders it. | Published. |
| /interop/attestation-corpus.json | The corpus of attestations gathered across rails. | Reachable only through a bulk API endpoint. | Published. |
| /interop/eas-attestation-batch.json | A prepared batch of attestation payloads for a third-party attestation service. | Reachable only through a bulk API endpoint. | Payloads are prepared. Nothing has been published to that service. |
| /interop/jailbreak-asr-evidence-pack.json | The per-model evidence behind the escape-detection work. | Signed, dated, and linked from nowhere. | Signed and published. |
| /interop/jail-peritem-v3.json | The item-by-item detail behind one board row, at the granularity a challenger needs. | Signed, dated, and linked from nowhere. | Signed and published. |
| /interop/mcp-security-scorecard.json | A security scorecard over tool servers, a measurement in its own right. | Reachable only through a bulk API endpoint. | Published. |
| /interop/card-store-verification.json | A record of searching five stores for the card bodies that would settle the disputed card count — and finding none. | This is a published negative result, which is exactly the kind of thing that should be easy to find and was impossible to find. | Published. The result was zero, and it is recorded as zero. |
| /interop/rwa-registry.json | The register of named financial instruments the financial rows draw from. | Referenced in the data layer, rendered on no page. | Published, and carrying no timestamp of any kind — so nothing derived from it can honestly claim a date. |
| /interop/index-reference-reverify.json | A re-check of the public reference series behind the candidate index rows. | Not listed in the estate's own surface catalogue, and linked from nowhere. | Published. |
| /interop/sbom-councilof-ai.json | The software bill of materials for this site. | Listed in the surface catalogue under a misspelt path, so a machine following the catalogue gets a 404. | Published at the corrected path. |
| /signed/chain.json | The full chain of card positions, including the ones whose contents are withheld. | The verification instructions tell a stranger to walk this chain, and no page linked it until now. | Published. |
Words used on this page
If a term appears above and is not explained here, that is a defect — tell us and we will either define it or stop using it.
- axis
- One thing we try to measure — a single question asked over and over, like “can this model tell which risk tier a system falls into?”. Every axis has its own set of questions and its own score.
- measured
- A real run happened: a fixed set of questions was asked, the answers were graded by a fixed rule, and the result was recorded and signed.
- unmeasured
- The slot is published, and nothing has been run against it. It appears on purpose so the gap is visible. It is never evidence that anything was measured.
- n
- How many things were actually measured. A score over ten items is a much weaker claim than the same score over three hundred, which is why n is always shown next to it.
- confidence range
- The band the true score is likely to sit in, given how few things were measured. A wide band means the headline number could easily move on a re-run. Two models whose bands overlap are not separated by the measurement.
- separated
- The leader's advantage over the rest of the field is large enough that it is unlikely to be chance. When it is not separated we call it a tie and refuse to publish an ordering.
- signed card
- A small file recording one measurement, stamped with a cryptographic signature. Anyone can re-check the stamp offline, without asking us and without trusting us.
- frozen bank
- The exact set of questions used, published and unchanged, so that anyone can ask a model the same questions and compare.