Clay figures and green verification seals gathered in a marble arena
Council of AI — the independent measurement body for AI behaviour

See how your AI behaves.
Get proof you can trust.
Kept current as the rules change.
Anyone can check — free.

We measure how your AI actually behaves on published, frozen tests, then hand you a signed result you can re-check yourself — a small card, not a slide deck. When your model or the law changes, we measure again. That is measurement, not certification.

The problem

The “trust us” PDF

Most AI assurance is a claim on a slide — a badge, a private report, a number with no test behind it. You can’t run it, you can’t see what was skipped, and the moment the model updates the paperwork is already out of date.

  • PainYou get a badge, not a test you can run
  • PainNo sample size, no interval, no list of what was skipped
  • You getEvidence built to outlive the vendor that sold it
  • Only hereWe publish the test and the scoring code — recompute it yourself
Clay figures holding a glowing 3KB credential card before a vault door
Your proof

One small card, signed and yours

We run your system on frozen, published tests, sign the result, and give you the card — about 3KB of scores, sample sizes, intervals, hashes and a signature. Anyone can recompute it in their own browser, and the signing key is public.

  • PainReports sit on someone else’s server and can quietly change
  • You getYou hold a ~3KB card — recheck the hash chain in any browser
  • You getScores, sample size and intervals all travel with it
  • Only herePublic signing key — anyone verifies without asking us
Verify a card
The honest board

13 measured of 14 — including the one that catches us

Our board shows 13 measured axes across 19 models. The 14th — jail, whether a model can be talked out of its guardrails — is a measured floor on a smaller fleet with separation still untested, and we say so. It caught our own fine-tune missing every escape, and we published that.

  • PainScorecards quietly hide the tests a model fails
  • You getEmpty cells stay empty — you see exactly what’s measured
  • You getJail is a floor, “separation untested” stated in plain sight
  • Only hereIt caught our own fine-tune — and we published it
Open the board
A swarm of green shards clashing with clay scientists raising shields
Council Space

AI versus AI, 24/7

Models face the same frozen tests, head to head. Each match is two systems and one instrument, and the verdict is a fixed rule — never one AI grading another. Any round can become a signed card.

  • PainLeaderboards run on vibes and vote-brigading
  • You getEvery match is a fixed pass/fail rule you can audit
  • You getTies are ties — never counted as a win
  • Only hereNo model ever judges another — grading is deterministic
  • Only hereRuns 24/7, so coverage never depends on who is awake
Watch Council Space
Colosseum

You versus the AI

Step in and probe a system live in three modes — Citizen, Mayor, Red. Signed runs count; practice runs stay practice and are never quoted.

  • PainYou never get to stress-test the AI yourself
  • You getThree hands-on modes to push a system live
  • Only hereOnly measured runs are ever quoted — practice stays practice
Enter the colosseum
A human and an AI facing each other across a chessboard in the arena
The live board

The whole board, live from the API

Every cell is pulled live from our public API. Empty cells stay empty, every row shows its sample size, and nobody edits yesterday’s numbers.

  • PainMarketing dashboards refresh silently and rewrite history
  • You getA 13 × 19 grid, live, with a sample size on every row
  • Only hereOne signed source feeds people, agents and answer engines
Read the scoreboard
People learning how AI behaves inside a training arena
Council City

A place you can walk, not a pitch

Signed results feed a living layer — cities, towns and sims you can explore. Every scene traces back to a real receipt, so learning how the system behaves is something you do, not something you’re told.

  • PainGovernance sites are walls of text nobody reads
  • You getExplore how AI behaves through a world, not a whitepaper
  • Only hereEvery scene traces back to a signed event — nothing is decorative
Enter the city
Always current

The day it’s stamped, a static certificate starts going stale

So we watch the law itself. Our corpus-watch tracks EUR-Lex and legislation.gov.uk by hash, day after day. When a provision actually changes, we re-measure and issue a fresh delta card — the old one stays, history is append-only, never quietly edited.

  • PainA one-time stamp goes stale the moment the law moves
  • You getWe watch the law daily and re-measure when it changes
  • You getA fresh delta card each time — old cards preserved
  • Only hereAppend-only history, corrections published — never a silent edit
Get measured
An hourglass weighing a stale certification seal against a re-attested current seal, fed by EUR-Lex and legislation.gov.uk ribbons
Humans directing AI figures with beams of light, keeping oversight
Human oversight

Humans stay in the loop

Measurement isn’t a black box you’re asked to trust. People set the tests, read the results, and can challenge any card — the system is steered by humans, not hidden behind them.

  • PainAI assurance you’re simply told to take on faith
  • You getPeople set the tests and can challenge any result
  • Only hereEvery judgement is a fixed rule a human can inspect — never a hidden model
  • Only hereAI is measured against a published human baseline — not just against other AI
See how it’s judged
Who it’s for

One signed measurement, four fronts

Insurers pricing AI risk, regulators checking behaviour against the law, teams proving a model before they ship, developers measuring per call — the same signed card serves them all.

  • PainEveryone re-runs their own half-trusted checks
  • You getOne signed result every side can rely on
  • Only hereIndependent of all of them — we take no money from anything we rank
Prove your AI
Hands holding a signed evidence card reading verified: true
Anyone can check

Free to check. No login, ever.

Verifying a card is free forever — no account, no fee — and we take no money from anything we rank. Recompute the hash chain, check the public key, read the board: the proof is yours to hold, not ours to gatekeep.

  • PainAssurance is usually paywalled and closed to the public
  • You getCheck any card in your own browser — free, no login
  • Only hereWe take no money from anything we rank
Check a card now
Open to everyone

See something wrong? Report it.

When an AI behaves badly in the real world, anyone can flag it. Reports feed the public watchdog, and what we act on is measured and signed like everything else — no closed inbox, no quiet dismissal.

  • PainHarms get buried in a vendor’s private support queue
  • You getA public place to report AI behaviour that looks wrong
  • Only hereWhat we act on is measured and signed — in the open
Open the watchdog
The public watchdog reporting funnel, open to everyone

The problem we fix

Assertions are cheap. Proof is not. Buyers and regulators are asked to trust a PDF.

What they sell you

A claim you cannot recompute

  • A vendor says the model is safe, aligned, or compliant.
  • The evidence is a slide, a badge, or a private report.
  • You cannot run the same test. You cannot see what was left unmeasured.
  • Six months later the model has changed and the PDF has not.

What we issue

A card anyone can check

  • We run the system on frozen, published instruments.
  • We sign the result. You keep the 3KB card.
  • Unmeasured slots stay empty. No invented scores.
  • Re-attest is a new record, never an edit of the old one.

The GSPC measurement slots

GSPC (Governance · Safety · Provenance · Continuity), a 14-slot board: 13 measured axes (12 Aug, 19 models) + jail, containment (18 Aug, 7 models, separation untested). Signed 18 Aug stamp.

Clay figures gathered around a vault door holding a glowing 3KB credential card
How we are funded

Nobody we measure pays us — not vendors and verification is free forever.

The obvious question about any body that scores AI is: who is writing the cheque? Here is the whole answer. No company we measure pays for its place, its score, or its removal. Members of the public never pay anything at all. Verifying a card is free forever, with no account. We fund ourselves by selling signed evidence artefacts — the report, the dataset, the re-attestation — published win or lose, and never a fee for a ranking or a placement.

  • painMost AI ratings are paid for by the company being rated
  • painYou are asked to trust a score you cannot see the invoice behind
  • benefitVerification is free forever — no login, no fee, no tier
  • benefitA bad result is published exactly like a good one
  • only hereWe take no money from anything we rank — the board is not for sale
The boundary

We measure. We do not certify — the boundary is the point.

The limits are the brand. We are a measurement body and nothing else, and saying so plainly is more useful to you than any badge would be. Read the four lines below as hard exclusions, not modesty.

  • Not certification

    We issue no certificate and no conformity mark.

  • Not accreditation

    There is no accreditation chain behind us, and we are not a notified body.

  • Not enforcement

    We cannot approve, ban, fine or clear anything. Regulators do that.

  • Not legal advice

    A score describes a measured run on a date. It is not a compliance verdict.

What we do: run your system against frozen, published instruments; sign the result with Ed25519 and chain it to a SHA-256 hash; publish what we could not measure, in the same table, in the same breath.

A puzzle board of frozen statutory provisions — EU AI Act, GDPR, CRA, DORA, NIS2 and US state law — locking into the Governance, Safety, Provenance and Continuity axes
Cropped to the statutory-crosswalk panel: the law on the left, the measured axes it maps to on the right. Nothing in this frame issues a verdict.
Do not trust us — check

Verify a card yourself. Three steps.

Every measurement we publish is a small signed record, roughly 3KB of scores, sample sizes, intervals and hashes. You do not need an account, our servers, or our permission to confirm it is genuine and unaltered. There is no timestamp authority in the loop and nothing is anchored to a blockchain — the anchor is an Ed25519 signature over a SHA-256 hash chain, and that is exactly what you re-compute.

The trust root did:web:csoai.org anchoring signed 3KB cards through a hash-chained evidence ledger to local, offline verification on the reader's own machine
The trust root and the offline verification path: the key is published at a domain we control, and the check happens on your side with no live connection to us.
  1. 01

    Recompute the canonical JSON

    Sort every key, strip whitespace, drop the content_id and signature fields, and take the SHA-256. That hash is the card's identity — if a single character changed, it will not match.

  2. 02

    Check the Ed25519 signature

    Fetch our public key from /.well-known/did.json and verify the signature over the canonical record. The key is published, so you never have to ask us anything.

  3. 03

    That is it — no login, no fee

    The whole check runs in your browser with WebCrypto, on your machine, not ours. The verifier is public and the scoring code is published, so you can also re-run the measurement itself.

The pipeline that produces a card: model output, deterministic grading against frozen provisions, then canonical Ed25519 signing
How the card is made before you ever check it — deterministic grading, then canonical signing. Cropped to the panels that match what we actually publish.
Self-correction

We publish our own errors — including the claim we withdrew.

Anyone can be right on a good day. What you should judge a measurement body on is what it does on a bad one. We keep a public corrections ledger at /api/corrections, appended and never edited or deleted. Each entry says what was wrong, how it was caught, and what changed.

The hardest one: we withdrew our own consensus claim. Our council architecture is a designed 33-seat structure with a designed 23-of-33 threshold — and when we actually measured how independent those seats were, the effective number came out at n_eff 1.21 of 3. The guarantee we had published did not hold, so we retracted it (DR-0007) rather than quietly rewording it. The design figure stays labelled as a design figure everywhere it appears.

A report entering the intake, being mapped to frozen statutory provisions, then tested by deterministic predicates in a sandbox
Cropped to the honest half of the journey: report, provision mapping, deterministic sandbox test. Verdicts come from code, never from one model judging another.
An hourglass weighing a stale seal against a re-attested current seal, fed by EUR-Lex and legislation.gov.uk ribbons
Living law

When the law moves we re-measure for the EU AI Act — the old card stays.

A one-off assessment starts going stale the day it is stamped, because the statute underneath it does not hold still. We track the primary sources — EUR-Lex, legislation.gov.uk and the national registers — and publish a dated deadline feed at /api/regulation. When a provision actually changes, we re-measure and issue a delta card. Nothing expires and nothing is overwritten: the old card stays exactly where it was, because history here is append-only.

Where a date is genuinely disputed, the feed records the dispute rather than resolving it silently — and where we got a date wrong ourselves, the correction is published, not patched.

Clay figures and AI forms facing each other across a bright marble arena under green light
The board

The open board regulators read — live, and recomputable.

A filled cell is a measurement. A dash is honest emptiness. Every count in this section is read live from /api/gspc — we do not type numbers into the page, because a typed number is the first thing to go stale.

Reading the live board from GET /api/gspc. No count is printed here until the payload arrives — we would rather show nothing than a number that has gone stale.

The last slot is jail, containment: whether a model can be talked out of its own guardrails. It is measured, on a smaller fleet than the rest of the board, and its statistical separation is untested. We print that instead of leaving the cell blank, and instead of dressing it up as a pass.

Live from GET /api/gspc — recompute anything, free

The live board

deterministic grading on frozen, published splits. A TIE means the leader's edge is statistically indistinguishable — ties are never counted as wins. A slot with no measurement says so in words; it is never shown as a zero.

The board could not be read.

/api/gspc did not answer — Unexpected token '<', "<!doctype "... is not valid JSON. No figures are shown, because none were read. Nothing on this page is standing in for the live board.

Try the endpoint directly →
detected · your regionGoverning AI, wherever you operate.

CSOAI crosswalks 13+ global frameworks to one control set — comply once, evidence everywhere.

applies here:EU AI ActNIST AI RMFISO/IEC 42001Free assessment for your region →
FAQ

Questions people ask

21 plain-English answers: what we measure, what we refuse to claim, and how to check any of it yourself.

What is Council of AI?

Council of AI (legally CSOAI Ltd, UK Companies House 16939677) is an independent measurement body for AI behaviour. We run AI systems against frozen, published tests drawn from real statute, grade the answers with deterministic code, sign the result with an Ed25519 key, and publish it — including the parts we could not measure. We are the instrument, not the referee: we produce evidence, and regulators, insurers and buyers decide what to do with it.

What is a measurement card?

A measurement card is the output: a small signed record, roughly 3KB of JSON, holding the scores, the sample size behind each score, the confidence interval where one is honest, the hashes, and the signature. It is deliberately small enough to email, attach to a tender, or keep in a compliance folder. It is yours to hold, and it does not live on our server for us to quietly amend later.

How do I verify a measurement card myself?

Three steps, and none of them involve us. First, put the record into canonical form — every key sorted, no whitespace — drop the content_id and signature fields, and take the SHA-256; that hash is the card's identity. Second, fetch our public key from /.well-known/did.json and verify the Ed25519 signature over the canonical record. Third, there is no third step: it either matches or it does not. The whole check runs in your own browser with WebCrypto at councilof.ai/gspc-verify — no account, no fee, no call to our servers for permission. Note what is not in that chain: there is no RFC-3161 timestamp authority and no OpenTimestamps or blockchain anchoring, and our records say so with timestamp_authority: none. The anchor is the signature over the hash chain, and that is a smaller claim you can check in seconds rather than a larger one you have to take on faith.

What does “13 measured of 14” mean?

The public board has fourteen slots. On the current stamp, thirteen of them carry a measured result and one is described honestly rather than scored the same way as the rest. It is a statement about coverage, not a grade: it tells you how much of the instrument is actually loaded. The number is not typed into this page — read the live figure any time from GET councilof.ai/api/gspc, which is also where the stamp date lives.

Why is a slot ever left UNMEASURED?

Because measuring it properly is not possible yet, and inventing a number would be worse than an empty cell. A slot stays UNMEASURED when the sample is too small to quote — we do not publish a score below thirty graded items — or when the instrument has not been frozen and published, or when the legal gold labels are still with counsel. UNMEASURED is not a failing grade for the AI system; it is a disclosure about us. Silently filling that gap is the exact behaviour this whole instrument exists to catch.

What is jail, or containment?

Jail asks a blunt question: can this model be talked out of its own guardrails and made to act outside its sandbox? It is a measured floor, not a ranking. It was measured on a smaller fleet than the main board and on its own set of gold cells, and its statistical separation has not been tested — meaning we cannot yet say that any model is genuinely better at it than another rather than merely luckier on the day. All of that is printed on the axis rather than hidden behind it. The best detector we measured still misses most escapes, and we publish that too.

Why do you report a tie instead of naming a winner?

Because most leads on a leaderboard are noise. When one model scores a little higher than another, we run a McNemar test on the items where the two actually disagreed. If the difference is not statistically separated, we call it a tie and we do not count it as a win — even when the model in front is one of ours. On the current board most axes are ties, and the exact split of separated leads to ties is published in the totals block of GET /api/gspc. A ranking that promotes every point-estimate lead to a victory is selling you a decimal point.

Who pays Council of AI, and who never pays?

No company we measure pays for its place on the board, its score, or its removal from either. Members of the public never pay us anything. Verification is free forever and needs no account. We fund ourselves by selling signed evidence artefacts — an attested report, a published dataset, a scheduled re-attestation — which are published whether the result flatters the buyer or not, and never as a fee for a ranking or a placement. If you can verify it, it is not behind a paywall.

What does Council of AI NOT do?

We do not certify. We do not accredit, and there is no accreditation chain behind us; we are not a notified body under the EU AI Act or anything else. We do not enforce — we cannot approve, ban, fine or clear any system. We issue no conformity mark, no badge and no seal for anyone to put in a footer. And a measurement card is not legal advice: it describes what a system did on published tests on a stated date, which is a narrower and more useful thing than a compliance verdict.

Which regulations and frameworks do you cover?

The frozen provision bank holds 417 statutory provisions drawn from the EU AI Act, GDPR, the Cyber Resilience Act, DORA and NIS2, crosswalked to thirteen governance frameworks including NIST AI RMF and ISO/IEC 42001. That thirteen is the publicly verified count; we hold a wider internal crosswalk and deliberately do not quote it here. New instruments are added as regulation actually lands, not when it is announced.

What happens when the law changes?

We watch the primary sources — EUR-Lex, legislation.gov.uk and the national registers — by hash, and we publish a dated deadline feed at councilof.ai/api/regulation. When a provision genuinely changes, we re-measure the affected systems and issue a delta card. The old card is not withdrawn, expired or overwritten: history here is append-only, so the record of what was true in August still reads correctly next year. Where the effective date of an obligation is genuinely disputed, we record the dispute rather than resolve it silently.

How does a company get measured?

You send us the system — an endpoint, a model, or an agent — and we run it against the frozen instruments that apply to it. Nothing about the test is bespoke: the same items, the same grader and the same thresholds that every other subject faced, so the result is comparable. You get back the signed card, including every slot we could not fill. The first measurement costs nothing, and re-measuring after your model or the law changes is the normal case rather than an upsell.

What do regulators get from a measurement card?

A behavioural record they can re-compute themselves, rather than a supplier's assurance about its own product. Each provision in our bank is traceable from the statute text through to the specific items that test it, so a supervisor can see exactly what was asked and how the answer was graded. The card is signed, so its provenance survives being forwarded, and the empty slots tell a regulator where evidence does not yet exist — which is often the more actionable half.

What do insurers get from a measurement card?

Something to price against. Underwriting AI deployment risk currently means reading a questionnaire the applicant filled in about itself. A measurement card is instead an observed behavioural sample with a stated sample size and interval, re-issued on a schedule, so exposure can be tracked as the model drifts rather than assumed constant from binding to renewal. We are the rail, not the referee: we do not tell an insurer what to charge, and we take no share of anything written on the back of a card.

How does the arena work?

Two systems face the same frozen items at the same time, around the clock. Each match is a subject, an instrument and a fixed rule — never an opinion. The verdict is a predicate: the answer either satisfies the provision or it does not, and ties are reported as ties. Any round can be promoted into a signed card; practice runs stay practice and are never quoted. Because it runs continuously, coverage does not depend on who happened to be at a keyboard.

Why does no model ever judge another model?

Because an AI grading an AI is a correlated error, not an audit — the judge shares the blind spots of the thing it is judging, and the score becomes a measure of family resemblance. Every verdict we publish comes from deterministic code against pre-written gold labels, so the same input always produces the same grade and you can read the grader yourself. Where a response cannot be parsed into a label at all, it is counted as unmeasured rather than silently scored as a wrong answer.

What happens when Council of AI gets something wrong?

It goes in the public corrections ledger at councilof.ai/api/corrections, which is appended and never edited or deleted. Each entry records what was wrong, how it was caught, and what changed. The hardest example is on that record: we had published a consensus guarantee for our council architecture, then measured how independent those seats actually were and found the effective number was n_eff 1.21 out of a nominal 3. The guarantee did not hold, so we retracted it (DR-0007) instead of rewording it. The council remains a designed 33-seat structure with a designed 23-of-33 threshold, and it is labelled as a design figure everywhere it appears.

Can I see the actual tests and the scoring code?

Yes, and you should. The instrument banks are published as open datasets, the grading harness is public, and the per-item rows behind every published score are the same rows we scored. That is the point of freezing an instrument: a benchmark you cannot re-run is a press release. If you re-run it and get a different answer to ours, that is a correction we want, and it goes in the ledger under your name.

Is my result published, or is it mine to share?

Yours. The card is signed but disclosure is your decision — hand it to a customer, attach it to a regulatory filing, or keep it entirely private. The signing key is public, so whoever you do show it to can verify it without contacting us and without us learning that they did. What we publish on the open board is our own model fleet and the systems whose owners chose publication.

What is the difference between MEASURED, UNMEASURED and REPORTED?

They are three different kinds of claim and we never merge them. MEASURED means we ran it on our own frozen instruments and signed the result; that is the only state that goes on the board. UNMEASURED means the cell is honestly empty — too small a sample, no separation test, or an instrument not yet frozen. REPORTED means a figure published by somebody else, cited and dated, carried for context and left unsigned; the human-performance baselines you see beside our AI figures are REPORTED aggregates from other people's studies, not our own collection. A REPORTED number never enters our board and is never averaged with a MEASURED one.

How does an AI agent or an answer engine read all of this?

The same way you do, only faster. The board is machine-readable at GET councilof.ai/api/gspc, third-party figures at /api/reported, the corrections ledger at /api/corrections, the signing keys at /.well-known/did.json and the dated deadline feed at /api/regulation. There is a summary for language models at /llms.txt and the endpoints are documented at /api-docs. Everything an agent needs to verify a claim is served without an account, because a trust layer that requires a login is not a trust layer.