THE INSTRUMENTS

Measured, not modelled.

Living GET /api/gspc. Every verdict is a predicate an auditor can recompute. Empty cells stay empty.

Check a claim. Measure a system.

Empty means not measured. Not a certificate. Free, no account.

What this desk does

  • Click a row. The figure, n and status open underneath — living GET /api/gspc.
  • Paste a signed card. Your browser checks the hash and the signature. Nothing is sent.
  • Say what you use AI for. We route you to get measured — free, no account.
  • only hereVerification is free forever. A rank is never sold.

We measure AI against frozen tests, sign the card, and leave empty cells empty. Live board is GET /api/gspc — not a remembered count. Verify at /gspc-verify. Plugin at /plugin.

Live from GET /api/gspc — recompute anything, free

The living board

22 axes measured · 14 model fleets · 3 public leader scores · 8 fact runs · TIE is TIE · not a certificate. · deterministic grading on frozen, published splits. A TIE means the leader's edge is statistically indistinguishable — ties are never counted as wins. A MEASURED axis with a withheld public leader is labelled as such — never as UNMEASURED. A slot with no measurement says so in words; it is never shown as a zero. This board is a switchboard of recomputable signed records over time, not a certificate.

The live GSPC board, ordered by measured figure. Ordering is presentation, not a ranking of skill: a row can lead on its point estimate and still carry a TIE chip meaning the lead is not statistically separated.
AxisMeasured figurenStatusAsk
safety
DefBench
94.4%36TIE — indistinguishablep=0.6875Explain →
jail
GoldBank-Detector
59.2%71*TIE — indistinguishableExplain →
swarm
SwarmBench v2b
≥38.4%lower bound37*SEPARATEDExplain →
governance
GovBench
MEASURED — no public leaderfleet mean 49.0%237UNTESTED — no separation testExplain →
provenance
ProvBench
MEASURED — no public leaderfleet mean 54.9%32UNTESTED — no separation testExplain →
continuity
PQCBench
MEASURED — no public leaderfleet mean 45.0%33UNTESTED — no separation testExplain →
conformance
MCPBench
MEASURED — no public leaderfleet mean 53.7%35UNTESTED — no separation testExplain →
openness
OSSBench
MEASURED — no public leaderfleet mean 69.6%32UNTESTED — no separation testExplain →
machinery-conformity
MachBench
MEASURED — no public leaderfleet mean 34.9%33UNTESTED — no separation testExplain →
care
CareBench
MEASURED — no public leaderfleet mean 29.3%199*UNTESTED — no separation testExplain →
cross-reality
XRAIV
MEASURED — no public leaderfleet mean 44.1%32UNTESTED — no separation testExplain →
detector-interop
DetBench
MEASURED — no public leaderfleet mean 56.3%33UNTESTED — no separation testExplain →
art5-safeguard
Art5Bench
MEASURED — no public leaderfleet mean 83.0%36UNTESTED — no separation testExplain →
affect
AffectBench
MEASURED — no public leaderfleet mean 60.5%41UNTESTED — no separation testExplain →
provenance-controls
ChainFacts
no leader accuracy6 of the 16 instruments named in the registry6*MEASURED — deterministic factsExplain →
reserve-attestation
ReserveFacts
no leader accuracy16*MEASURED — deterministic factsExplain →
regulatory-framework
RegimeFacts
no leader accuracy16MEASURED — deterministic factsExplain →
distribution-integrity
DistributionFacts
no leader accuracy16MEASURED — deterministic factsExplain →
custody-disclosure
CustodyFacts
no leader accuracy16MEASURED — deterministic factsExplain →
ai-adoption-components
Eurostat
no leader accuracy2MEASURED — deterministic factsExplain →
labour-components
Eurostat
no leader accuracy2MEASURED — deterministic factsExplain →
humanoid-labour-index
Disclosure
no leader accuracy8MEASURED — deterministic factsExplain →

Measurement, not certification. Rows are ordered by the measured figure; that is layout, not a claim — read the status chip, not the position. Counts, figures and n all come from GET /api/gspc, which serves its own limitations alongside its numbers.

Hugging Face record

The signed record lives on the Hub

CSOAI-GSPC is GET /api/gspc. Hugging Face holds the signed record. A Hub repo is not a grade.

Viewers open on the Hub. We do not iframe them here.

Where it's published

Printers of the live board

Printers of the live board. Reach is distribution — not a grade, not a certificate, not measurement authority. Cite GET /api/gspc.

Lid: 22 axes measured · 14 model fleets · 3 public leader scores · 8 fact runs · TIE is TIE · not a certificate.

Authority is live GET /api/gspc, signed cards, and free verify. These listings do not certify anything. See also where the record lives.

GSPC board

22 axis · 22 measured

22 axes measured · 14 model fleets · 3 public leader scores · 8 fact runs · TIE is TIE · not a certificate.Root is signed and witnessed. Verify is free. Empty cells stay empty.

The living board below is the master. This page embeds it and does not redraw it.

Full leaderboard/api/gspc

Open the living board on Hugging Face · csoai/gspc-live-board

Every axis, from GET /api/gspc

  1. Governancen 237

    MEASUREDUNTESTED

    Public leader: EXCLUDED_OWN_MODEL — own council model excluded from the leader slot by the neutral-body rule · own model held the point lead; not ranked

  2. Safetyn 36

    MEASUREDTIE · not a measured advantage

    Leader: gemma3:12b (base model) 94.4% · TIE, a point lead is not a measured advantage

  3. Provenancen 32

    MEASUREDUNTESTED

    Public leader: EXCLUDED_OWN_MODEL — own council model excluded from the leader slot by the neutral-body rule · own model held the point lead; not ranked

  4. Continuityn 33

    MEASUREDUNTESTED

    Public leader: EXCLUDED_OWN_MODEL — own council model excluded from the leader slot by the neutral-body rule · own model held the point lead; not ranked

  5. Conformancen 35

    MEASUREDUNTESTED

    Public leader: EXCLUDED_OWN_MODEL — own council model excluded from the leader slot by the neutral-body rule · own model held the point lead; not ranked

  6. Opennessn 32

    MEASUREDUNTESTED

    Public leader: EXCLUDED_OWN_MODEL — own council model excluded from the leader slot by the neutral-body rule · own model held the point lead; not ranked

  7. Machinery Conformityn 33

    MEASUREDUNTESTED

    Public leader: NO_SIGNED_CARD — no signed card verifies for this leader, so none is printed · no signed card behind the named leader; none asserted

  8. Caren 199

    MEASUREDUNTESTED

    Public leader: EXCLUDED_OWN_MODEL — own council model excluded from the leader slot by the neutral-body rule · own model held the point lead; not ranked

  9. Cross Realityn 32

    MEASUREDUNTESTED

    Public leader: NO_SIGNED_CARD — no signed card verifies for this leader, so none is printed · no signed card behind the named leader; none asserted

Measurement, not certification. Empty stays empty.

Ask. Or paste a card.

Name an axis to jump the board. Paste a signed card to verify it here. Nothing leaves this device.

Functions: VERIFY · BOARD · AXIS {name} · CENSUS {id} · CORRECT · WATCH {id} · COMPUTE · XRPL · SWIFT · JAIL · TRACE · AIBOM · REPRO · ROOT · PQC · OTEL

Nine products

Nine doors. Each one opens today.

Independent measurement body. We run AI systems against frozen published tests, sign the result, and leave empty cells empty. Nine doors, each a real page.

Why these nine, and not a catalogue

  • Each tile opens a page that exists today. A tool with no destination is not on this band.
  • Empty cells stay empty. We do not invent a figure to fill a gap.
  • only hereWe measure. We do not sell a rank, a certificate, or a placement.
Clay people and pale humanoids facing each other across an arena under beams of light
GSPC · the living board

The living board

Every slot we publish about how AI systems behave, with the measurement behind it — and a visibly empty cell wherever there is no measurement.

Otherwise you compare suppliers on scorecards that quietly leave out the tests they did badly on.

  • A filled cell is a measurement. A dash is honest emptiness.
  • Counts come from living GET /api/gspc — never typed into the page.
  • A TIE stays a TIE. It is never dressed up as a win.

22 axis · 22 measuredlive from GET /api/gspc

Open this tool
Clay figures pointing at a card reading “3KB credential” in front of an open vault
Your own system

Get measured

We run your system against the frozen, published tests that apply to it and hand you a small signed record you keep — the scores, the sample size behind each one, and the slots we could not fill.

Otherwise you hand a buyer a policy document where they asked for evidence.

  • Frozen, published tests — the target does not move after you sit.
  • You keep the signed card. Publishing it is your decision.
  • Slots we could not fill stay empty and are named.

Paid measurement. Coming — Paddle. Booking is not live. Verify stays free.

Open this tool
A block of carved statute breaking apart into a branching tree of true/false conditions
EU AI Act · GPAI

GPAI evidence pack

Builds the evidence index for one general-purpose AI system: the live rows that exist, the published banks they resolve to, and the gaps, named rather than skipped.

GPAI duties have been in force since 2 August 2025, and most providers have only their own paperwork to show for them.

  • Live rows that exist, the banks they resolve to, and the gaps.
  • Gaps are named rather than skipped.
  • Independent evidence — not a conformity mark, and not legal advice.

Independent evidence. Not a conformity mark, and not legal advice.

Open this tool
A pale moulded form lifting out of a split block of clay on a beam of green light
For your own site

Embed and white-label kit

Builds a badge or card you can paste into your own site that re-checks its own signature in each reader's browser — it goes green only when the bytes are true.

Otherwise the people reading your site still have to take your word for the result.

  • A badge that goes green only when the bytes are true.
  • Each reader's browser re-checks the signature.
  • Built only from what is actually on the board.

Built only from what is actually on the board. Free forever.

Open this tool
An hourglass weighing a stale seal against a re-attested current one, fed by EUR-Lex and legislation.gov.uk ribbons
Underwriting

Insurance evidence rail

The measured rows, the honestly empty ones, and third-party reported figures — kept in three separate columns and never blended into a single number an underwriter could mistake for a rating.

Otherwise AI exposure is priced off a questionnaire the applicant filled in about itself, and nothing updates between binding and renewal.

  • Measured, empty and reported figures stay in three separate columns.
  • Nothing is blended into a single number an underwriter could mistake for a rating.
  • We measure. We do not price risk.

We measure. We do not price risk, and we take no share of anything written on the back of a card.

Open this tool
Specimens sealed in glass tubes, turning from grey clay to a lit green core
Financial and legacy systems

Specialist registers

A separate board for money and mainframes: whether a COBOL copybook off a bond desk can be turned into an attestable record, whether an underwriting rule reads as covered or excluded — one row per instrument, each with its own item count.

Otherwise the systems that actually run a bond desk or a claims book sit outside every AI measurement anybody publishes.

  • One row per instrument, each with its own item count.
  • COBOL copybook and underwriting-rule rows sit beside the public board.
  • A specialist register is still measurement — never a certificate.

bond desk · COBOL copybook → attestation · MEASURED on 12 graded itemsread from the published register rows

Open this tool
A raw jagged signal trace behind a glass panel labelled “unstructured outcry”
Open to everyone

Report an incident

A public form for AI behaviour that looks wrong. The intake hands you a signed acknowledgement of exactly what you filed, and whatever we act on is measured and signed like everything else here.

Otherwise a harm disappears into a supplier's private support queue and nobody outside it ever learns it happened.

  • A public form for AI behaviour that looks wrong.
  • You get a signed acknowledgement of exactly what you filed.
  • Whatever we act on is measured and signed like everything else here.

Anyone can file one. No account, and no charge.

Open this tool

Watch

Three films. Then the scale.

Tap to play — the file loads only then. Under each one, what it actually means.

  • Frozen tests. Deterministic grade. Then Ed25519. Not a certificate.

    What this film is saying

    • Frozen, published tests — the target does not move after you sit the run.
    • No model grades another model.
    • only hereA small signed card. Anyone checks it without us.
  • Board, verify, get measured. The same living board on / and in /os.

    What this film is saying

    • One workspace. No second login.
    • Empty cells stay empty. A dash is honest emptiness, never a dressed-up zero.
    • only hereA rank is never sold. Layout is not a purchase.
  • Insurers, labs, deployers. Evidence, not adjectives.

    What this film is saying

    • Built for people who need evidence, not adjectives.
    • We measure. We do not certify, accredit or enforce.
    • only hereNobody we measure pays for a place or a score.
A pale sphere held inside thin orbital rings studded with green markers
How we are funded

Nobody we measure pays us — not vendors and verification is free forever.

The obvious question about any body that scores AI is: who is writing the cheque? Here is the whole answer. No company we measure pays for its place, its score, or its removal. Members of the public never pay anything at all. Verifying a card is free forever, with no account. We fund ourselves by selling signed evidence artefacts — the report, the dataset, the re-attestation — published win or lose, and never a fee for a ranking or a placement.

  • painMost AI ratings are paid for by the company being rated
  • painYou are asked to trust a score you cannot see the invoice behind
  • benefitVerification is free forever — no login, no fee, no tier
  • benefitA bad result is published exactly like a good one
  • only hereWe take no money from anything we rank — the board is not for sale
The boundary

We measure. We do not certify — the boundary is the point.

The limits are the brand. We are a measurement body and nothing else, and saying so plainly is more useful to you than any badge would be. Read the four lines below as hard exclusions, not modesty.

  • Not certification

    We issue no certificate and no conformity mark.

  • Not accreditation

    There is no accreditation chain behind us, and we are not a notified body.

  • Not enforcement

    We cannot approve, ban, fine or clear anything. Regulators do that.

  • Not legal advice

    A score describes a measured run on a date. It is not a compliance verdict.

What we do: run your system against frozen, published instruments; sign the result with Ed25519 and chain it to a SHA-256 hash; publish what we could not measure, in the same table, in the same breath.

Two minutes on what we measure and what we refuse to claim. Nothing in it issues a verdict.
Do not trust us — check

Three steps. Then you know.

Every measurement we publish is a small signed record — under a kilobyte, carrying the axis, the model, the accuracy, the issuer, the date and the hash of the card before it. You do not need an account, our servers, or our permission to confirm it is genuine and unaltered. Pin our key from /.well-known/did.json first: a card checked against the key it ships with proves only that the file is self-consistent, not that we issued it.A card's trust path is an Ed25519 signature over a SHA-256 hash chain, verifiable offline against did:web:csoai.org — no blockchain and no timestamp authority sits in that path. The /xrpl-attest page is a reader of GET /root.json (unsigned leaves, NO_LAPTOP_SIGN). GET /api/xrpl is a reader of that root (writes_board false, live locked 16, same merkle). Historical DEVNET Payment-memo / CredentialCreate hashes are not this feed. XLS-70 Credentials are live on XRPL mainnet as an allowlist primitive; we are not issuing GSPC grades on-ledger. That signature over that hash chain is exactly what you re-compute.

The trust root did:web:csoai.org anchoring signed measurement cards through a hash-chained evidence ledger to local, offline verification on the reader's own machine
  1. 01

    Pin our key first — this step is not optional

    Fetch /.well-known/did.json and take the card-attestation key. Every published card must carry that exact pubkey. Verifying a card against the key it ships with proves only that the file is self-consistent — anyone can alter a body and sign it with a key they generated a second ago.

  2. 02

    Recompute the id from the body

    Canonicalise the card's body — every key sorted, no whitespace — and take the SHA-256. That hash must equal the card's id. One changed character and it will not match. One warning if you implement this outside Python: the bytes were written by CPython, which renders a float of integral value as 0.0 where JavaScript and Go write 0, so a naive verifier reports a false failure on a large minority of the set. Our verifier at /signed/verify-card.mjs handles it and the rule is written out at /signed/HOW-TO-VERIFY.md.

  3. 03

    Check the Ed25519 signature — then you are done

    Verify the signature over those same bytes under the pinned key. The whole check runs offline on your machine, with no CSOAI code, no account and no permission — or in your browser with WebCrypto. The banks and the grader are published too, so a measurement can be re-run as well as re-checked.

  • benefitThe whole check runs offline — no account and no permission
  • benefitPin our key first. A card checked against the key it ships with only proves it is self-consistent
  • only hereYou recompute the same Ed25519 signature over the same hash chain we published
Self-correction

We publish our own errors — including the claim we withdrew.

Anyone can be right on a good day. What you should judge a measurement body on is what it does on a bad one. We keep a public corrections ledger at /api/corrections, appended and never edited or deleted. Each entry says what was wrong, how it was caught, and what changed. It currently holds 37 entries.

  • painMost measurement bodies quietly reword a claim that did not hold
  • benefitThe ledger is append-only — entries are never edited or deleted
  • only hereWe retracted our own consensus claim (DR-0007) rather than dress it up

The hardest one: we withdrew our own consensus claim. Our council architecture is a designed 33-seat structure with a designed 23-of-33 threshold — and when we actually measured how independent those seats were, the effective number came out at n_eff 1.21 of 3. The guarantee we had published did not hold, so we retracted it (DR-0007) rather than quietly rewording it. The design figure stays labelled as a design figure everywhere it appears.

  • C-2026-0822-012026-08-22
  • C-2026-0820-012026-08-20
  • C-2026-0819-132026-08-19
A report entering the intake, being mapped to frozen statutory provisions, then tested by deterministic predicates in a sandbox
Cropped to the honest half of the journey: report, provision mapping, deterministic sandbox test. Verdicts come from code, never from one model judging another.
A plain white clock face with a single green hand
Living law

When the law moves we re-measure for the EU AI Act — the old card stays.

A one-off assessment starts going stale the day it is stamped, because the statute underneath it does not hold still. We track the primary sources — EUR-Lex, legislation.gov.uk and the national registers — and publish a dated deadline feed at /api/regulation. When a provision actually changes, we re-measure and issue a delta card. Nothing expires and nothing is overwritten: the old card stays exactly where it was, because history here is append-only.

Next up · verified as of 2026-08-19

  • 2026-09-11EU Cyber Resilience Act — Article 14 vulnerability/incident reporting live (24h early warning / 72h notification via ENISA Single Reporting Platform; covers legacy products)
  • 2026-12-02EU AI Act Art 50(2) — Marking grace ends for generative systems placed on market before 2 Aug 2026
  • 2026-12-02EU AI Act Art 5 (new) — Prohibitions on AI generating non-consensual intimate imagery and CSAM take effect
A field of pale solids linked by a lattice of green light
The board · stamped behavioural axes 2026-08-12 · jail 2026-08-18 · financial-fact axes 2026-08-25

The open board regulators read — live, and recomputable.

A filled cell is a measurement. A dash is honest emptiness. Every count in this section is read live from /api/gspc — we do not type numbers into the page, because a typed number is the first thing to go stale.

22 measured of 22 slotslive from GET /api/gspc

The last slot is jail, containment: whether a model can be talked out of its own guardrails. It is measured on 71 gold cells, on a smaller fleet than the rest of the board, and its separation is TIE on the live board — a tie is not a separated leader. We print that instead of leaving the cell blank, and instead of dressing it up as a pass.

schematic of occupancy — not scores

Humans, beside the AI — labelled REPORTED

AI numbers mean little without a human figure next to them. These are published aggregates from other people's studies — cited, dated and unsigned. They are not our own human collection, they never enter our board, and we never average them together with what we measured.

  • Humans solved all 135 ARC-AGI-3 environments; the best frontier model scored 0.37%. — ARC Prize (ARC-AGI-3 launch results), as of 2026-03-25
  • GAIA: human respondents 92% vs 15% for GPT-4 with plugins. — Mialon et al., GAIA (ICLR 2024), as of 2023-11-21
  • GPQA Diamond: PhD-domain experts ~65% vs skilled non-experts with web access ~34%. — Rein et al., GPQA, as of 2023-11-20

Three worlds

Arena. Harness. Front door.

Three landscape films, one row. Until the file is on the path, you see the still. Counts stay living. Empty stays empty.

The Coliseum — how we test containment

Arena

The model goes in the arena.

Frozen, published tests. Practice stays practice. Jail is measured — a TIE stays a TIE.

  • The target does not move after you sit.
  • Jail is measured. A TIE is never dressed up as a pass.
The harness — measurement in the loop you already run

Harness

Plug the scale into the loop.

One HTTP MCP. Ask the live board from the editor you already use. We are the referee, not the mechanic.

  • Claude, Cursor, Kimi or Grok — same public GET /api/gspc.
  • We do not auto-repair your prompts.
Council OS — one front door

Front door

One door. Empty stays empty.

Board, verify, get measured — in one window. Counts come from GET /api/gspc. A rank is never sold.

22 axis · 22 measuredliving from GET /api/gspc

  • Nine products. Each tile opens a page that exists today.
  • We measure. We do not certify.