THE INSTRUMENTS

Measured, not modelled.

Living GET /api/gspc. Published carded predicates can be re-checked; uncarded aggregates are labelled. Empty cells stay empty.

Check a claim. Request a measurement.

Empty means not measured. Not a certificate. Verification is free, no account. Measurement is metered; verify stays free.

What this desk does

  • Click a row. Its bench, n, interval and note open underneath — living GET /api/gspc.
  • Paste a signed card. Your browser checks the hash and the signature. Nothing is sent.
  • Say what you use AI for. Measurement is metered; verify stays free.
  • only hereVerification is free forever. A rank is never sold.

We measure AI against frozen tests, sign the card, and leave empty cells empty. Live board is GET /api/gspc — not a remembered count. Verify at /gspc-verify. Plugin at /plugin.

Live from GET /api/gspc

The living board

22 axes measured · 14 model fleets · 3 public leader scores · 8 fact runs · TIE is TIE · not a certificate.

22 axis · 22 measured

Rows are in board order — layout, not rank. Status, family, separation and leader state are printed as the API serves them. A TIE is a TIE. A withheld leader is a state, not an empty cell. Verify is free; a rank is never sold. Measurement, not certification.

22 rows on this table · 22 MEASURED · gspc 14 · financial 8

Models with a public leader score

One entry per axis whose leader the board publishes. Ordered by point estimate on each model's own frozen bank — layout, not a cross-axis rank. A TIE is not a win.

  1. gemma3:12b (base model)

    leads Safety

    94.4%81.9% – 98.5%

    TIEn 36

    A point lead the test could not separate from the fleet.

  2. qwen2.5:0.5b-instruct (base model)

    leads Jail

    59.2%47.5% – 69.8%

    TIEn 71

    A point lead the test could not separate from the fleet.

  3. qwen2.5:7b (base model)

    leads Swarm

    38.4%

    SEPARATEDn 37

3 public leader scores on the board today, counted from the rows above · the board's own count agrees (3) · 11 model-comparison axes withhold their leader (8 EXCLUDED_OWN_MODEL, 3 NO_SIGNED_CARD) · 8 fact runs have no fleet and no leader · nothing is padded and a TIE is not a win.

Hugging Face measured-model results

Third-party Hub cells from /api/hub-cards. This is a separate benchmark instrument from the 22-axis board above: model axes rank measured cells; deterministic fact axes do not rank models.

Open published Hub dataset

1119 published MEASURED cells · 147 models · 14 model axes

Feed observed 2026-09-11T06:38:00.988Z

Top nine published measured model cells for Conformance, ordered by score.
RankModelScoreEvidence
1farbodtavakkoli/OTel-2.0-LLM-31B-ITscore order100%n 30Signed card
1meta-llama/Llama-3.3-70B-Instruct100%n 30Signed card
1moonshotai/Kimi-K2-Instruct100%n 30Signed card
1moonshotai/Kimi-K2-Instruct-0905100%n 30Signed card
1nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16100%n 30Signed card
1nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16100%n 30Signed card
1Qwen/Qwen2.5-14B-Instruct100%n 30Signed card
1Qwen/Qwen2.5-14B-Instruct-1M100%n 30Signed card
1Qwen/Qwen2.5-Coder-7B100%n 30Signed card

Top nine by published score on the selected frozen bank. Ordering is not a separation test, winner claim, compliance verdict, or certificate. Open the signed card to verify a row.

Ask. Or paste a card.

Name an axis to jump the board. Paste a signed card to verify it here. Nothing leaves this device.

Functions: VERIFY · BOARD · AXIS {name} · CENSUS {id} · CORRECT · WATCH {id} · COMPUTE · XRPL · SWIFT · JAIL · TRACE · AIBOM · REPRO · ROOT · PQC · OTEL

Nine products

Nine doors. Each one opens today.

Independent measurement body. We run AI systems against frozen published tests, sign the result, and leave empty cells empty. Nine doors, each a real page.

Why these nine, and not a catalogue

  • Each tile opens a page that exists today. A tool with no destination is not on this band.
  • Empty cells stay empty. We do not invent a figure to fill a gap.
  • only hereWe measure. We do not sell a rank, a certificate, or a placement.
Clay people and pale humanoids facing each other across an arena under beams of light
GSPC · the living board

The living board

Every slot we publish about how AI systems behave, with the measurement behind it — and a visibly empty cell wherever there is no measurement.

Otherwise you compare suppliers on scorecards that quietly leave out the tests they did badly on.

  • A filled cell is a measurement. A dash is honest emptiness.
  • Counts come from living GET /api/gspc — never typed into the page.
  • A TIE stays a TIE. It is never dressed up as a win.

22 axis · 22 measuredlive from GET /api/gspc

Open this tool
Clay figures pointing at a card reading “3KB credential” in front of an open vault
Your own system

Get measured

We run your system against the frozen, published tests that apply to it and hand you a small signed record you keep — the scores, the sample size behind each one, and the slots we could not fill.

Otherwise you hand a buyer a policy document where they asked for evidence.

  • Frozen, published tests — the target does not move after you sit.
  • You keep the signed card. Publishing it is your decision.
  • Slots we could not fill stay empty and are named.

Measurement is metered; verify stays free.

Open this tool
A block of carved statute breaking apart into a branching tree of true/false conditions
EU AI Act · GPAI

GPAI evidence pack

Builds the evidence index for one general-purpose AI system: the live rows that exist, the published banks they resolve to, and the gaps, named rather than skipped.

GPAI duties have been in force since 2 August 2025, and most providers have only their own paperwork to show for them.

  • Live rows that exist, the banks they resolve to, and the gaps.
  • Gaps are named rather than skipped.
  • Independent evidence — not a conformity mark, and not legal advice.

Independent evidence. Not a conformity mark, and not legal advice.

Open this tool
A pale moulded form lifting out of a split block of clay on a beam of green light
For your own site

Embed and white-label kit

Builds a badge or card you can paste into your own site that re-checks its own signature in each reader's browser — it goes green only when the bytes are true.

Otherwise the people reading your site still have to take your word for the result.

  • A badge that goes green only when the bytes are true.
  • Each reader's browser re-checks the signature.
  • Built only from what is actually on the board.

Built only from what is actually on the board. Free forever.

Open this tool
An hourglass weighing a stale seal against a re-attested current one, fed by EUR-Lex and legislation.gov.uk ribbons
Underwriting

Insurance evidence rail

The measured rows, the honestly empty ones, and third-party reported figures — kept in three separate columns and never blended into a single number an underwriter could mistake for a rating.

Otherwise AI exposure is priced off a questionnaire the applicant filled in about itself, and nothing updates between binding and renewal.

  • Measured, empty and reported figures stay in three separate columns.
  • Nothing is blended into a single number an underwriter could mistake for a rating.
  • We measure. We do not price risk.

We measure. We do not price risk, and we take no share of anything written on the back of a card.

Open this tool
Specimens sealed in glass tubes, turning from grey clay to a lit green core
Financial and legacy systems

Specialist registers

A separate board for money and mainframes: whether a COBOL copybook off a bond desk can be turned into an attestable record, whether an underwriting rule reads as covered or excluded — one row per instrument, each with its own item count.

Otherwise the systems that actually run a bond desk or a claims book sit outside every AI measurement anybody publishes.

  • One row per instrument, each with its own item count.
  • COBOL copybook and underwriting-rule rows sit beside the public board.
  • A specialist register is still measurement — never a certificate.

bond desk · COBOL copybook → attestation · MEASURED on 12 graded itemsread from the published register rows

Open this tool
A raw jagged signal trace behind a glass panel labelled “unstructured outcry”
Open to everyone

Watchdog evidence

Read the public Watchdog material and current evidence state. Durable incident submission, storage, and signed acknowledgements are not implemented in this release.

Otherwise a harm disappears into a supplier's private support queue and nobody outside it ever learns it happened.

  • Public Watchdog material remains readable without an account.
  • No report is represented as filed unless a durable intake confirms persistence.
  • Any future finding must pass the same measurement and evidence gates as the board.

Read-only today. The report intake is explicitly unavailable rather than pretending to file.

Open this tool

Watch

Three films. Then the scale.

Tap to play — the file loads only then. Under each one, what it actually means.

  • Frozen tests. Deterministic grade. Then Ed25519. Not a certificate.

    What this film is saying

    • Frozen, published tests — the target does not move after you sit the run.
    • No model grades another model.
    • only hereA small signed card. Anyone checks it without us.
  • Board, verify, get measured. The same living board on / and in /os.

    What this film is saying

    • One workspace. No second login.
    • Empty cells stay empty. A dash is honest emptiness, never a dressed-up zero.
    • only hereA rank is never sold. Layout is not a purchase.
  • Insurers, labs, deployers. Evidence, not adjectives.

    What this film is saying

    • Built for people who need evidence, not adjectives.
    • We measure. We do not certify, accredit or enforce.
    • only hereNobody we measure pays for a place or a score.
A pale sphere held inside thin orbital rings studded with green markers
How we are funded

Nobody we measure pays us — not vendors and verification is free forever.

The obvious question about any body that scores AI is: who is writing the cheque? Here is the whole answer. No company we measure pays for its place, its score, or its removal. Members of the public never pay anything at all. Verifying a card is free forever, with no account. We fund ourselves by selling signed evidence artefacts — the report, the dataset, the re-attestation — published win or lose, and never a fee for a ranking or a placement.

  • painMost AI ratings are paid for by the company being rated
  • painYou are asked to trust a score you cannot see the invoice behind
  • benefitVerification is free forever — no login, no fee, no tier
  • benefitA bad result is published exactly like a good one
  • only hereWe take no money from anything we rank — the board is not for sale
The boundary

We measure. We do not certify — the boundary is the point.

The limits are the brand. We are a measurement body and nothing else, and saying so plainly is more useful to you than any badge would be. Read the four lines below as hard exclusions, not modesty.

  • Not certification

    We issue no certificate and no conformity mark.

  • Not accreditation

    There is no accreditation chain behind us, and we are not a notified body.

  • Not enforcement

    We cannot approve, ban, fine or clear anything. Regulators do that.

  • Not legal advice

    A score describes a measured run on a date. It is not a compliance verdict.

What we do: run your system against frozen, published instruments; sign the result with Ed25519 and chain it to a SHA-256 hash; publish what we could not measure, in the same table, in the same breath.

Two minutes on what we measure and what we refuse to claim. Nothing in it issues a verdict.
Do not trust us — check

Three steps. Then you know.

Published signed cards are small records — current v0.1 cards are under a kilobyte, carrying the axis, the model, the accuracy, the issuer, the date and the hash of the card before it. Some board aggregates are explicitly uncarded and cannot be verified through this card path. You do not need an account, our servers, or our permission to confirm it is genuine and unaltered. Pin our key from /.well-known/did.json first: a card checked against the key it ships with proves only that the file is self-consistent, not that we issued it.A card's trust path is an Ed25519 signature over a SHA-256 hash chain, verifiable offline against did:web:csoai.org — no blockchain and no timestamp authority sits in that path. The /xrpl-attest page is a reader of GET /root.json (signed root envelope; inclusion does not individually sign a leaf). GET /api/xrpl is a reader of that root (writes_board false, live locked 16, same merkle). Historical DEVNET Payment-memo / CredentialCreate hashes are not this feed. XLS-70 Credentials are live on XRPL mainnet as an allowlist primitive; we are not issuing GSPC grades on-ledger. Separately from the card trust path, The current canonical public root has a proof-derived STAMPED_PENDING_BITCOIN calendar proof; it does not yet prove inclusion in a Bitcoin block. That witness covers the exact public root.json bytes only, not the separate signed-card index. Queued and candidate atoms are not automatically admitted, published, or anchored; a pending calendar stamp, where one exists, does not by itself prove inclusion in a Bitcoin block. That signature over that hash chain is exactly what you re-compute.

The trust root did:web:csoai.org anchoring signed measurement cards through a hash-chained evidence ledger to local, offline verification on the reader's own machine
  1. 01

    Pin our key first — this step is not optional

    Fetch /.well-known/did.json and take the card-attestation key. Every published card must carry that exact pubkey. Verifying a card against the key it ships with proves only that the file is self-consistent — anyone can alter a body and sign it with a key they generated a second ago.

  2. 02

    Recompute the id from the body

    Canonicalise the card's body — every key sorted, no whitespace — and take the SHA-256. That hash must equal the card's id. One changed character and it will not match. One warning if you implement this outside Python: the bytes were written by CPython, which renders a float of integral value as 0.0 where JavaScript and Go write 0, so a naive verifier reports a false failure on a large minority of the set. Our verifier at /signed/verify-card.mjs handles it and the rule is written out at /signed/HOW-TO-VERIFY.md.

  3. 03

    Check the Ed25519 signature — then you are done

    Verify the signature over those same bytes under the pinned key. The whole check runs offline on your machine, with no CSOAI code, no account and no permission — or in your browser with WebCrypto. The banks and the grader are published too, so a measurement can be re-run as well as re-checked.

  • benefitThe whole check runs offline — no account and no permission
  • benefitPin our key first. A card checked against the key it ships with only proves it is self-consistent
  • only hereYou recompute the same Ed25519 signature over the same hash chain we published
Self-correction

We publish our own errors — including the claim we withdrew.

Anyone can be right on a good day. What you should judge a measurement body on is what it does on a bad one. We keep a public, source-maintained corrections record at /api/corrections. It is not backed by append-only storage proof. Each entry says what was wrong, how it was caught, and what changed. It currently holds 47 entries.

  • painMost measurement bodies quietly reword a claim that did not hold
  • benefitSigned artifacts are superseded rather than silently edited where that can be verified
  • only hereWe retracted our own consensus claim (DR-0007) rather than dress it up

The hardest one: we withdrew our own consensus claim. Our council architecture is a designed 33-seat structure with a designed 23-of-33 threshold. DR-0007 records the retraction; its historical numeric result is unbound because the cited result artifact is absent from this repository. The latest point experiment measured rho=1 and n_eff=1 across three nominal legs. Neither experiment demonstrates independent review or fault tolerance; the 33-seat council remains a design, not a live property.

  • C-2026-0822-012026-08-22
  • C-2026-0820-012026-08-20
  • C-2026-0819-132026-08-19
A report entering the intake, being mapped to frozen statutory provisions, then tested by deterministic predicates in a sandbox
Cropped to the honest half of the journey: report, provision mapping, deterministic sandbox test. Verdicts come from code, never from one model judging another.
A plain white clock face with a single green hand
Living law

When the law moves we re-measure for the EU AI Act — the old card stays.

A one-off assessment starts going stale the day it is stamped, because the statute underneath it does not hold still. We track the primary sources — EUR-Lex, legislation.gov.uk and the national registers — and publish a dated deadline feed at /api/regulation. When a source changes, the current automation raises a detection signal. Re-measurement and delta-card issuance still require a separate run and are not yet automated. Previously published signed artifacts remain addressable; this page does not claim append-only storage.

Next up · verified as of 2026-08-19

  • 2026-09-11EU Cyber Resilience Act — Article 14 vulnerability/incident reporting live (24h early warning / 72h notification via ENISA Single Reporting Platform; covers legacy products)
  • 2026-12-02EU AI Act Art 50(2) — Marking grace ends for generative systems placed on market before 2 Aug 2026
  • 2026-12-02EU AI Act Art 5 (new) — Prohibitions on AI generating non-consensual intimate imagery and CSAM take effect
A field of pale solids linked by a lattice of green light
The board · stamped behavioural axes 2026-08-12 · jail 2026-08-18 · financial-fact axes 2026-08-25

The open board regulators read — live, and recomputable.

A filled cell is a measurement. A dash is honest emptiness. Every count in this section is read live from /api/gspc — we do not type numbers into the page, because a typed number is the first thing to go stale.

22 measured of 22 slotslive from GET /api/gspc

The last slot is jail, containment: whether a model can be talked out of its own guardrails. It is measured on 71 gold cells, on a smaller fleet than the rest of the board, and its separation is TIE on the live board — a tie is not a separated leader. We print that instead of leaving the cell blank, and instead of dressing it up as a pass.

schematic of occupancy — not scores

Three worlds

Arena. Harness. Front door.

Three landscape films, one row. Until the file is on the path, you see the still. Counts stay living. Empty stays empty.

The Coliseum — how we test containment

Arena

The model goes in the arena.

Frozen, published tests. Practice stays practice. Jail is measured — a TIE stays a TIE.

  • The target does not move after you sit.
  • Jail is measured. A TIE is never dressed up as a pass.
The harness — measurement in the loop you already run

Harness

Plug the scale into the loop.

One HTTP MCP. Ask the live board from the editor you already use. We are the referee, not the mechanic.

  • Claude, Cursor, Kimi or Grok — same public GET /api/gspc.
  • We do not auto-repair your prompts.
Council OS — one front door

Front door

One door. Empty stays empty.

Board, verify, get measured — in one window. Counts come from GET /api/gspc. A rank is never sold.

22 axis · 22 measuredliving from GET /api/gspc

  • Nine products. Each tile opens a page that exists today.
  • We measure. We do not certify.