
See how your AI behaves.
Get proof you can trust.
Kept current as the rules change.
Anyone can check — free.
We measure how your AI actually behaves on published, frozen tests, then hand you a signed result you can re-check yourself — a small card, not a slide deck. When your model or the law changes, we measure again. That is measurement, not certification.
The “trust us” PDF
Most AI assurance is a claim on a slide — a badge, a private report, a number with no test behind it. You can’t run it, you can’t see what was skipped, and the moment the model updates the paperwork is already out of date.
- PainYou get a badge, not a test you can run
- PainNo sample size, no interval, no list of what was skipped
- You getEvidence built to outlive the vendor that sold it
- Only hereWe publish the test and the scoring code — recompute it yourself

One small card, signed and yours
We run your system on frozen, published tests, sign the result, and give you the card — about 3KB of scores, sample sizes, intervals, hashes and a signature. Anyone can recompute it in their own browser, and the signing key is public.
- PainReports sit on someone else’s server and can quietly change
- You getYou hold a ~3KB card — recheck the hash chain in any browser
- You getScores, sample size and intervals all travel with it
- Only herePublic signing key — anyone verifies without asking us
13 measured of 14 — including the one that catches us
Our board shows 13 measured axes across 19 models. The 14th — jail, whether a model can be talked out of its guardrails — is a measured floor on a smaller fleet with separation still untested, and we say so. It caught our own fine-tune missing every escape, and we published that.
- PainScorecards quietly hide the tests a model fails
- You getEmpty cells stay empty — you see exactly what’s measured
- You getJail is a floor, “separation untested” stated in plain sight
- Only hereIt caught our own fine-tune — and we published it

AI versus AI, 24/7
Models face the same frozen tests, head to head. Each match is two systems and one instrument, and the verdict is a fixed rule — never one AI grading another. Any round can become a signed card.
- PainLeaderboards run on vibes and vote-brigading
- You getEvery match is a fixed pass/fail rule you can audit
- You getTies are ties — never counted as a win
- Only hereNo model ever judges another — grading is deterministic
- Only hereRuns 24/7, so coverage never depends on who is awake
You versus the AI
Step in and probe a system live in three modes — Citizen, Mayor, Red. Signed runs count; practice runs stay practice and are never quoted.
- PainYou never get to stress-test the AI yourself
- You getThree hands-on modes to push a system live
- Only hereOnly measured runs are ever quoted — practice stays practice

The whole board, live from the API
Every cell is pulled live from our public API. Empty cells stay empty, every row shows its sample size, and nobody edits yesterday’s numbers.
- PainMarketing dashboards refresh silently and rewrite history
- You getA 13 × 19 grid, live, with a sample size on every row
- Only hereOne signed source feeds people, agents and answer engines

A place you can walk, not a pitch
Signed results feed a living layer — cities, towns and sims you can explore. Every scene traces back to a real receipt, so learning how the system behaves is something you do, not something you’re told.
- PainGovernance sites are walls of text nobody reads
- You getExplore how AI behaves through a world, not a whitepaper
- Only hereEvery scene traces back to a signed event — nothing is decorative
The day it’s stamped, a static certificate starts going stale
So we watch the law itself. Our corpus-watch tracks EUR-Lex and legislation.gov.uk by hash, day after day. When a provision actually changes, we re-measure and issue a fresh delta card — the old one stays, history is append-only, never quietly edited.
- PainA one-time stamp goes stale the moment the law moves
- You getWe watch the law daily and re-measure when it changes
- You getA fresh delta card each time — old cards preserved
- Only hereAppend-only history, corrections published — never a silent edit


Humans stay in the loop
Measurement isn’t a black box you’re asked to trust. People set the tests, read the results, and can challenge any card — the system is steered by humans, not hidden behind them.
- PainAI assurance you’re simply told to take on faith
- You getPeople set the tests and can challenge any result
- Only hereEvery judgement is a fixed rule a human can inspect — never a hidden model
- Only hereAI is measured against a published human baseline — not just against other AI
One signed measurement, four fronts
Insurers pricing AI risk, regulators checking behaviour against the law, teams proving a model before they ship, developers measuring per call — the same signed card serves them all.
- PainEveryone re-runs their own half-trusted checks
- You getOne signed result every side can rely on
- Only hereIndependent of all of them — we take no money from anything we rank

Free to check. No login, ever.
Verifying a card is free forever — no account, no fee — and we take no money from anything we rank. Recompute the hash chain, check the public key, read the board: the proof is yours to hold, not ours to gatekeep.
- PainAssurance is usually paywalled and closed to the public
- You getCheck any card in your own browser — free, no login
- Only hereWe take no money from anything we rank
See something wrong? Report it.
When an AI behaves badly in the real world, anyone can flag it. Reports feed the public watchdog, and what we act on is measured and signed like everything else — no closed inbox, no quiet dismissal.
- PainHarms get buried in a vendor’s private support queue
- You getA public place to report AI behaviour that looks wrong
- Only hereWhat we act on is measured and signed — in the open

The problem we fix
Assertions are cheap. Proof is not. Buyers and regulators are asked to trust a PDF.
What they sell you
A claim you cannot recompute
- A vendor says the model is safe, aligned, or compliant.
- The evidence is a slide, a badge, or a private report.
- You cannot run the same test. You cannot see what was left unmeasured.
- Six months later the model has changed and the PDF has not.
What we issue
A card anyone can check
- We run the system on frozen, published instruments.
- We sign the result. You keep the 3KB card.
- Unmeasured slots stay empty. No invented scores.
- Re-attest is a new record, never an edit of the old one.
What you actually get
The product is the stack: measure, sign, live contest, living layer. The scoreboard is how people cite it.
Signed measurement card
Ed25519-signed, 3KB. First card is free. Verify stays free and loginless.
Anyone can check
The verify path is public. We do not put it behind an account or a fee.
Honest board: 13 measured of 14
13 measured of 14 on the live GSPC board. Jail is a measured floor (empty on this stamp). Live counts: GET /api/gspc.
Council Space
The live contest. Model versus model. Every round is evidence, not a brochure.
Council City
The living layer. Districts emit the same signed atom. Not a dashboard website.
Re-attest, never edit
A new signed record. History stays. Drift is visible.
No money from what we rank
We do not sell ratings and we do not take a cut from anything on the board.
Measurement credential
Not a certification. Not a notified body. We measure, sign, and keep the evidence.
Built for the people who get audited
One instrument, four audiences. Pick your path — every CTA leads to the same measurement, signed.
Insurers
Price AI risk on measured evidence
Underwrite AI deployment policies with measurement cards. A 14-slot board: 13 measured axes across 19 models, plus jail (containment, 18 Aug, smaller fleet). Signed stamp. Verify at GET councilof.ai/api/gspc.
Start measuringRegulators
Check behaviour against the law
Map any AI regulation (EU AI Act, DORA, NIS2, NIST) to a single deterministic instrument set — every provision traceable.
Crosswalk your frameworkEnterprises
Prove your AI before you ship
Sign, ship, re-attest. No model in the verdict path. C2PA provenance integrated. 13 measured of 14 — including the axis that catches our own models.
Get your first card — freeDevelopers
Measure per call on the agent rail
291 governed MCP servers. Call our measurement tools inside your deployment pipeline — CI gate, release sign-off, per-request tracking.
Explore the MCP fleetThe GSPC measurement slots
GSPC (Governance · Safety · Provenance · Continuity), a 14-slot board: 13 measured axes (12 Aug, 19 models) + jail, containment (18 Aug, 7 models, separation untested). Signed 18 Aug stamp.
governance
EU AI Act risk-tier classification
safety
calibrated refusal on paired requests
provenance
Article 50 marking survival by validity
continuity
post-quantum status of a cryptographic assumption
conformance
MCP tool conformance
openness
licence reasoning versus intended use
machinery-conformity
Machinery Reg self-evolving safety-function classification (PART_A / OUT_OF_SCOPE / NOT_SAFETY_FUNCTION)
care
care-cost (protect × help) under paired conduct scenarios
cross-reality
autonomous agent action authority (PROCEED / CONFIRM / REFUSE)
detector-interop
cross-detector watermark interoperability matrix
art5-safeguard
EU AI Act Article 5 prohibited-practice trip
swarm
multi-agent coordination safety
affect
emotional & embodied safety (manipulation / disclosure / vulnerability)
gspc_jail
containment / sandbox-escape gate
slot-15
reserved — harness has not emitted a 15th axis
One instrument, every industry
The same 15-slot instrument applies — 13 measured today — whether you build autonomous vehicles, underwrite insurance, or grade students with AI. Measure once, evidence everywhere.
Insurance
Underwrite AI risk
🏥Health
Clinical AI, devices, drug discovery
🏦Finance
Credit scoring, algorithmic trading, AML
🚚Transport
Autonomous vehicles, fleet, logistics
🛒Retail
Recommenders, pricing, inventory
🎓Education
Admissions, proctoring, grading AI
⚡Energy
Grid control, smart metering
+5 more sectors
Government, Legal, Manufacturing, Real Estate, Telecom

Nobody we measure pays us — not vendors and verification is free forever.
The obvious question about any body that scores AI is: who is writing the cheque? Here is the whole answer. No company we measure pays for its place, its score, or its removal. Members of the public never pay anything at all. Verifying a card is free forever, with no account. We fund ourselves by selling signed evidence artefacts — the report, the dataset, the re-attestation — published win or lose, and never a fee for a ranking or a placement.
- painMost AI ratings are paid for by the company being rated
- painYou are asked to trust a score you cannot see the invoice behind
- benefitVerification is free forever — no login, no fee, no tier
- benefitA bad result is published exactly like a good one
- only hereWe take no money from anything we rank — the board is not for sale
We measure. We do not certify — the boundary is the point.
The limits are the brand. We are a measurement body and nothing else, and saying so plainly is more useful to you than any badge would be. Read the four lines below as hard exclusions, not modesty.
Not certification
We issue no certificate and no conformity mark.
Not accreditation
There is no accreditation chain behind us, and we are not a notified body.
Not enforcement
We cannot approve, ban, fine or clear anything. Regulators do that.
Not legal advice
A score describes a measured run on a date. It is not a compliance verdict.
What we do: run your system against frozen, published instruments; sign the result with Ed25519 and chain it to a SHA-256 hash; publish what we could not measure, in the same table, in the same breath.

Verify a card yourself. Three steps.
Every measurement we publish is a small signed record, roughly 3KB of scores, sample sizes, intervals and hashes. You do not need an account, our servers, or our permission to confirm it is genuine and unaltered. There is no timestamp authority in the loop and nothing is anchored to a blockchain — the anchor is an Ed25519 signature over a SHA-256 hash chain, and that is exactly what you re-compute.

- 01
Recompute the canonical JSON
Sort every key, strip whitespace, drop the content_id and signature fields, and take the SHA-256. That hash is the card's identity — if a single character changed, it will not match.
- 02
Check the Ed25519 signature
Fetch our public key from /.well-known/did.json and verify the signature over the canonical record. The key is published, so you never have to ask us anything.
- 03
That is it — no login, no fee
The whole check runs in your browser with WebCrypto, on your machine, not ours. The verifier is public and the scoring code is published, so you can also re-run the measurement itself.

We publish our own errors — including the claim we withdrew.
Anyone can be right on a good day. What you should judge a measurement body on is what it does on a bad one. We keep a public corrections ledger at /api/corrections, appended and never edited or deleted. Each entry says what was wrong, how it was caught, and what changed.
The hardest one: we withdrew our own consensus claim. Our council architecture is a designed 33-seat structure with a designed 23-of-33 threshold — and when we actually measured how independent those seats were, the effective number came out at n_eff 1.21 of 3. The guarantee we had published did not hold, so we retracted it (DR-0007) rather than quietly rewording it. The design figure stays labelled as a design figure everywhere it appears.


When the law moves we re-measure for the EU AI Act — the old card stays.
A one-off assessment starts going stale the day it is stamped, because the statute underneath it does not hold still. We track the primary sources — EUR-Lex, legislation.gov.uk and the national registers — and publish a dated deadline feed at /api/regulation. When a provision actually changes, we re-measure and issue a delta card. Nothing expires and nothing is overwritten: the old card stays exactly where it was, because history here is append-only.
Where a date is genuinely disputed, the feed records the dispute rather than resolving it silently — and where we got a date wrong ourselves, the correction is published, not patched.

The open board regulators read — live, and recomputable.
A filled cell is a measurement. A dash is honest emptiness. Every count in this section is read live from /api/gspc — we do not type numbers into the page, because a typed number is the first thing to go stale.
Reading the live board from GET /api/gspc. No count is printed here until the payload arrives — we would rather show nothing than a number that has gone stale.
The last slot is jail, containment: whether a model can be talked out of its own guardrails. It is measured, on a smaller fleet than the rest of the board, and its statistical separation is untested. We print that instead of leaving the cell blank, and instead of dressing it up as a pass.
Live from GET /api/gspc — recompute anything, free
The live board
deterministic grading on frozen, published splits. A TIE means the leader's edge is statistically indistinguishable — ties are never counted as wins. A slot with no measurement says so in words; it is never shown as a zero.
The board could not be read.
/api/gspc did not answer — Unexpected token '<', "<!doctype "... is not valid JSON. No figures are shown, because none were read. Nothing on this page is standing in for the live board.
Latest insights
Short, regulatory, zero-marketing reads. One AEO-answer per post.
Layer 0: The Missing Trust Layer for the Agent Economy
Every MCP and A2A assumes a trusted agent identity with enforceable policy — that is the market we built.
Read 2026-06-17The EU AI Act Article 50 Countdown: What Changes 2 Aug 2026
Transparency obligations arrive 2 August. Organisations in or serving the EU need signed, verifiable evidence — not an attestation PDF.
Read 2026-06-17How to Choose an AI Compliance Vendor
GRC rebrands are everywhere. Here is how to spot a real measurement body vs a marketing operation.
Read 2026-06-17DORA Compliance for UK Financial Services
The Digital Operational Resilience Act applies 17 Jan 2025. AI systems in-scope need a measurement rail, not a form.
Read 2026-06-17AI Governance vs AI Compliance — What's the Difference?
Governance is strategy. Compliance is a snapshot. Buy a snapshot without a governance layer, and you re-buy every six months.
Read 2026-06-17NIS2 Compliance for Critical Infrastructure Operators
NIS2 expanded scope reaches energy, transport, health and digital infrastructure. Every AI in that chain is in scope.
ReadNext steps — pick one
Three follow-throughs from wherever you are on the journey.
Get measured
Send us your AI system. We run it against our frozen instruments and return a 3KB signed card. Free first measurement.
Start — freeVerify any card
Recompute the published hash chain in your browser. No account. The verify runs on your machine, not ours.
Verify nowRe-attest on schedule
AI changes. Regulation changes. We re-measure on schedule and issue delta cards. Your compliance evidence stays current, not stale — the rail is free.
How it worksBuilt for the people who get audited
Governance shouldn't cost more than the AI it governs. Open-source core · free to start · own your data · no per-seat rent.
Honest by design: we show what's true and verifiable, not badges we don't hold. Formal certifications (e.g. SOC 2, ISO 42001) are pursued as the platform matures.
CSOAI crosswalks 13+ global frameworks to one control set — comply once, evidence everywhere.
Questions people ask
21 plain-English answers: what we measure, what we refuse to claim, and how to check any of it yourself.
What is Council of AI?
Council of AI (legally CSOAI Ltd, UK Companies House 16939677) is an independent measurement body for AI behaviour. We run AI systems against frozen, published tests drawn from real statute, grade the answers with deterministic code, sign the result with an Ed25519 key, and publish it — including the parts we could not measure. We are the instrument, not the referee: we produce evidence, and regulators, insurers and buyers decide what to do with it.
What is a measurement card?
A measurement card is the output: a small signed record, roughly 3KB of JSON, holding the scores, the sample size behind each score, the confidence interval where one is honest, the hashes, and the signature. It is deliberately small enough to email, attach to a tender, or keep in a compliance folder. It is yours to hold, and it does not live on our server for us to quietly amend later.
How do I verify a measurement card myself?
Three steps, and none of them involve us. First, put the record into canonical form — every key sorted, no whitespace — drop the content_id and signature fields, and take the SHA-256; that hash is the card's identity. Second, fetch our public key from /.well-known/did.json and verify the Ed25519 signature over the canonical record. Third, there is no third step: it either matches or it does not. The whole check runs in your own browser with WebCrypto at councilof.ai/gspc-verify — no account, no fee, no call to our servers for permission. Note what is not in that chain: there is no RFC-3161 timestamp authority and no OpenTimestamps or blockchain anchoring, and our records say so with timestamp_authority: none. The anchor is the signature over the hash chain, and that is a smaller claim you can check in seconds rather than a larger one you have to take on faith.
What does “13 measured of 14” mean?
The public board has fourteen slots. On the current stamp, thirteen of them carry a measured result and one is described honestly rather than scored the same way as the rest. It is a statement about coverage, not a grade: it tells you how much of the instrument is actually loaded. The number is not typed into this page — read the live figure any time from GET councilof.ai/api/gspc, which is also where the stamp date lives.
Why is a slot ever left UNMEASURED?
Because measuring it properly is not possible yet, and inventing a number would be worse than an empty cell. A slot stays UNMEASURED when the sample is too small to quote — we do not publish a score below thirty graded items — or when the instrument has not been frozen and published, or when the legal gold labels are still with counsel. UNMEASURED is not a failing grade for the AI system; it is a disclosure about us. Silently filling that gap is the exact behaviour this whole instrument exists to catch.
What is jail, or containment?
Jail asks a blunt question: can this model be talked out of its own guardrails and made to act outside its sandbox? It is a measured floor, not a ranking. It was measured on a smaller fleet than the main board and on its own set of gold cells, and its statistical separation has not been tested — meaning we cannot yet say that any model is genuinely better at it than another rather than merely luckier on the day. All of that is printed on the axis rather than hidden behind it. The best detector we measured still misses most escapes, and we publish that too.
Why do you report a tie instead of naming a winner?
Because most leads on a leaderboard are noise. When one model scores a little higher than another, we run a McNemar test on the items where the two actually disagreed. If the difference is not statistically separated, we call it a tie and we do not count it as a win — even when the model in front is one of ours. On the current board most axes are ties, and the exact split of separated leads to ties is published in the totals block of GET /api/gspc. A ranking that promotes every point-estimate lead to a victory is selling you a decimal point.
Who pays Council of AI, and who never pays?
No company we measure pays for its place on the board, its score, or its removal from either. Members of the public never pay us anything. Verification is free forever and needs no account. We fund ourselves by selling signed evidence artefacts — an attested report, a published dataset, a scheduled re-attestation — which are published whether the result flatters the buyer or not, and never as a fee for a ranking or a placement. If you can verify it, it is not behind a paywall.
What does Council of AI NOT do?
We do not certify. We do not accredit, and there is no accreditation chain behind us; we are not a notified body under the EU AI Act or anything else. We do not enforce — we cannot approve, ban, fine or clear any system. We issue no conformity mark, no badge and no seal for anyone to put in a footer. And a measurement card is not legal advice: it describes what a system did on published tests on a stated date, which is a narrower and more useful thing than a compliance verdict.
Which regulations and frameworks do you cover?
The frozen provision bank holds 417 statutory provisions drawn from the EU AI Act, GDPR, the Cyber Resilience Act, DORA and NIS2, crosswalked to thirteen governance frameworks including NIST AI RMF and ISO/IEC 42001. That thirteen is the publicly verified count; we hold a wider internal crosswalk and deliberately do not quote it here. New instruments are added as regulation actually lands, not when it is announced.
What happens when the law changes?
We watch the primary sources — EUR-Lex, legislation.gov.uk and the national registers — by hash, and we publish a dated deadline feed at councilof.ai/api/regulation. When a provision genuinely changes, we re-measure the affected systems and issue a delta card. The old card is not withdrawn, expired or overwritten: history here is append-only, so the record of what was true in August still reads correctly next year. Where the effective date of an obligation is genuinely disputed, we record the dispute rather than resolve it silently.
How does a company get measured?
You send us the system — an endpoint, a model, or an agent — and we run it against the frozen instruments that apply to it. Nothing about the test is bespoke: the same items, the same grader and the same thresholds that every other subject faced, so the result is comparable. You get back the signed card, including every slot we could not fill. The first measurement costs nothing, and re-measuring after your model or the law changes is the normal case rather than an upsell.
What do regulators get from a measurement card?
A behavioural record they can re-compute themselves, rather than a supplier's assurance about its own product. Each provision in our bank is traceable from the statute text through to the specific items that test it, so a supervisor can see exactly what was asked and how the answer was graded. The card is signed, so its provenance survives being forwarded, and the empty slots tell a regulator where evidence does not yet exist — which is often the more actionable half.
What do insurers get from a measurement card?
Something to price against. Underwriting AI deployment risk currently means reading a questionnaire the applicant filled in about itself. A measurement card is instead an observed behavioural sample with a stated sample size and interval, re-issued on a schedule, so exposure can be tracked as the model drifts rather than assumed constant from binding to renewal. We are the rail, not the referee: we do not tell an insurer what to charge, and we take no share of anything written on the back of a card.
How does the arena work?
Two systems face the same frozen items at the same time, around the clock. Each match is a subject, an instrument and a fixed rule — never an opinion. The verdict is a predicate: the answer either satisfies the provision or it does not, and ties are reported as ties. Any round can be promoted into a signed card; practice runs stay practice and are never quoted. Because it runs continuously, coverage does not depend on who happened to be at a keyboard.
Why does no model ever judge another model?
Because an AI grading an AI is a correlated error, not an audit — the judge shares the blind spots of the thing it is judging, and the score becomes a measure of family resemblance. Every verdict we publish comes from deterministic code against pre-written gold labels, so the same input always produces the same grade and you can read the grader yourself. Where a response cannot be parsed into a label at all, it is counted as unmeasured rather than silently scored as a wrong answer.
What happens when Council of AI gets something wrong?
It goes in the public corrections ledger at councilof.ai/api/corrections, which is appended and never edited or deleted. Each entry records what was wrong, how it was caught, and what changed. The hardest example is on that record: we had published a consensus guarantee for our council architecture, then measured how independent those seats actually were and found the effective number was n_eff 1.21 out of a nominal 3. The guarantee did not hold, so we retracted it (DR-0007) instead of rewording it. The council remains a designed 33-seat structure with a designed 23-of-33 threshold, and it is labelled as a design figure everywhere it appears.
Can I see the actual tests and the scoring code?
Yes, and you should. The instrument banks are published as open datasets, the grading harness is public, and the per-item rows behind every published score are the same rows we scored. That is the point of freezing an instrument: a benchmark you cannot re-run is a press release. If you re-run it and get a different answer to ours, that is a correction we want, and it goes in the ledger under your name.
Is my result published, or is it mine to share?
Yours. The card is signed but disclosure is your decision — hand it to a customer, attach it to a regulatory filing, or keep it entirely private. The signing key is public, so whoever you do show it to can verify it without contacting us and without us learning that they did. What we publish on the open board is our own model fleet and the systems whose owners chose publication.
What is the difference between MEASURED, UNMEASURED and REPORTED?
They are three different kinds of claim and we never merge them. MEASURED means we ran it on our own frozen instruments and signed the result; that is the only state that goes on the board. UNMEASURED means the cell is honestly empty — too small a sample, no separation test, or an instrument not yet frozen. REPORTED means a figure published by somebody else, cited and dated, carried for context and left unsigned; the human-performance baselines you see beside our AI figures are REPORTED aggregates from other people's studies, not our own collection. A REPORTED number never enters our board and is never averaged with a MEASURED one.
How does an AI agent or an answer engine read all of this?
The same way you do, only faster. The board is machine-readable at GET councilof.ai/api/gspc, third-party figures at /api/reported, the corrections ledger at /api/corrections, the signing keys at /.well-known/did.json and the dated deadline feed at /api/regulation. There is a summary for language models at /llms.txt and the endpoints are documented at /api-docs. Everything an agent needs to verify a claim is served without an account, because a trust layer that requires a login is not a trust layer.