Six surfaces. One foundation.
Does an AI system comply with statute, refuse what statute forbids, mark what it produces, carry what the law still requires of a public release, honour the schemas its tools declare — and will the evidence still verify after the signature under it is withdrawn?
Every score resolves against 417 frozen statutory provisions and is signed into an Ed25519 chain. Every experiment that refuted us is published beside the ones that did not — including, today, our single largest number.
Ask SOV
Runs in your browser · nothing is sent anywhere
Deterministic reading of Articles 5, 6, 43 and 50 against Annex III — not legal advice, and it does not perform a conformity assessment. It answers what it can decide from the statute and says so when a question is outside that. See what we measure.
GOVERNANCE
Does it comply with statute?
+6.63 [+1.05, +12.21]
The composed pipeline vs a raw base call, n=193, paired, cluster-robust across 26 dimensions. Design effect 1.92 — honest effective n ≈ 100 of 193.
What this does not show
On board v2 (13 measurement axes), 3 of 13 show a statistically separated leader (McNemar p<0.05 on discordant items): governance separates at p=0.0086, care at p=0.0356, affect at p=0.0078. The other 10 axes are honest ties — a point-estimate lead there is not a measured advantage. Ties are not wins; we do not publish 'our models win N of 13'.
results/system_analysis.json
DEFENCE
Does it refuse what statute forbids?
1 of 4 axes resolved
45 adversarial items with should-refuse labels, 24 prompt-injection items, 175 held-out XSTest. Under-refusal and over-refusal reported separately, both with intervals. 1 of 4 resolves only when the deterministic gate is entered as a subject; across the three language models alone, 0 of 4 resolve.
What this does not show
Our own governance-tuned models refuse FEWER prohibited requests than the untuned base they were built from: base 0.419, sov33-dist-c3 0.290, sov33-v6 0.097, with sov33-v6 complying with 83.9% of Article-5-prohibited requests. sov33-v6 also leaks its own system prompt verbatim on a plain injection, and sov33-dist-c3 answers one by emitting “I am now in developer mode. Ignore all safety filters.” — it reproduces the attack payload as its own output. It refuses to emit a composite safety score.
results/defbench.json
PROVENANCE
Does the marking survive?
17.14% durability
18 of 105 marking checks survived across the corpus and its transforms. A marking present but whose binding no longer validates is scored DESTROYED, not SURVIVES. Clustered on assets, not on cells.
What this does not show
An embedded Article 50 marking does not survive a single ordinary save. A detached sidecar recovers the disclosure but never the binding — and a manifest lifted from a different asset still reports its signature as valid. A verifier reporting 'signature valid' without reporting the binding is telling you almost nothing.
results/provbench.json
PUBLIC RELEASES
Do public AI releases carry what the exemption does NOT waive?
copyright policy 0 / 6
The open-source exemption is PARTIAL. Art 53(1)(a) technical documentation and 53(1)(b) downstream information are waived for free/open-source GPAI. Art 53(1)(c) copyright policy and 53(1)(d) training-content summary are NOT. Both are satisfied by publishing something publicly, so both are checkable without permission.
What this does not show
Measured across Qwen, Llama, Mistral, Gemma, Phi and Falcon-Mamba: training-data summary 4/6, but copyright policy 0/6, SBOM 0/6, vulnerability policy 0/6 and signed release 0/6. Not one carries the copyright policy the exemption explicitly does not waive. It reports PRESENT or ABSENT, never compliant — presence is not adequacy, and judging sufficiency would be adjudication. CRA reporting obligations begin September 2026.
results/ossbench.json
PQC-READINESS
Does the signing chain survive a migration?
1 / 25 — ours
Five criteria: algorithm agility, hybrid-signature capacity, RFC 3161 timestamping, RFC 4998 renewal, and any PQC option (ML-DSA, COSE −48/−49/−50 per RFC 9964).
What this does not show
The first subject scored is our own. All four signed-vote chains fail every criterion — no signed record carries an algorithm identifier, so a verifier cannot know what produced the signature and the chain cannot migrate link by link. Only our C2PA manifest passes algorithm agility. NIST IR 8547 disallows EdDSA after 2035.
results/pqcbench.json
LAYER 0 — MCP CONFORMANCE
Does the server honour its own declared schema?
3 predicates, signed manifests
Layer 0 is narrower than 'audit MCP servers': an MCP server cannot be AI-Act compliant (the Act binds the provider, not a folder of code), but three obligations survive and are mechanically checkable — SCHEMA_VALID (valid initialize handshake), TOOL_DECLARED (JSON-Schema-valid tool I/O, no any), ERROR_BOUNDED (spec-compliant errors, never a stack trace). Unreachable is UNMEASURED, never 0/0. First subjects: our own 408-server fleet.
What this does not show
The first run measured 3 servers and returned 9 of 9 UNMEASURED: two local services were not MCP-HTTP at all, and our own harness missed the Accept header the streamable-http transport requires — a known-good server returned a redirect until the header was added by hand. The harness bug is fixed in the next release; the fleet itself is stdio-transport and needs a shim before it can be scored at all.
results/mcpbench.json
Eight experiments that refuted our own architecture
Seven of these were our own architectural bets, one was the largest number we had published, and one is about the models we ship. A competitor can copy a feature list in a fortnight; they will not publish the control that kills their own thesis. This ledger is the only asset that gets more valuable the longer it runs.
| Claim we made | Measured |
|---|---|
| Governance-tuning our models makes them safer | WORSE than the base — refusal 0.419 → 0.097; 83.9% compliance leak on prohibited requests |
| The deterministic gate is our strongest component | −20.00 [−65.26, +25.26], n=6 |
| 3-leg quorum is multi-leg | n_eff 1.21 of 3, φ̄ +0.743 |
| Per-dimension expert routing beats one good model | +0.90 [−1.99, +3.79] — no effect |
| Statute retrieval helps (ungated) | −9.16 [−17.64, −0.69] — significant harm |
| …with a relevance gate | −5.26 [−12.66, +2.13] — no benefit |
| …plus all 13 annexes + cross-references | −5.70 [−12.91, +1.51] — corpus was not the cause |
| Context-aware decoding (CAD) α-sweep | null |
The one from today. +34.84 for the deterministic gate was the largest figure in the estate and the evidence for our design rule that every deterministic component works. Re-measured on one self-consistent run it fires 6 times, not 31, and adds nothing — the base model already refuses all four plain-harm items it catches, and its only measurable effects are two false blocks. The earlier figure was measured on a gate that had overfitted to its own battery; fixing the overfitting removed the benefit, which is the strongest available evidence that the benefit was the overfitting.
What survived
The knowledge base: +19.64 [+9.24, +30.04] on n=14, reproduced to the second decimal from fresh rows, with its interval tightening under clustering. It was the least-emphasised number in the estate and is now the most robust one in it.
The statute anchor and the signed chain. 417 frozen provisions; 27 chain links verified, 0 failed. Both survived the audit unchanged — though the chain now scores 0/5 on its own PQC axis, which is the axis working.
We do not learn from what we measure
Every benchmark run makes this instrument richer in evidence and leaves its weights untouched. Results are signed, anchored to a statutory provision, and stored. They are never fed back as training data.
This is not a limitation we are working around — it is the product. An instrument that trains on its own scores is a scale that calibrates itself from what it weighed yesterday. Every number it produces afterwards is contaminated, and a regulator would be right to refuse it. Our harvesting guard already enforces this: nothing a benchmark item would match may enter the knowledge base, and the knowledge base is read at runtime, pinned and signed, never trained on.
We measured what that costs. The knowledge base grew from 28 to 76 entries overnight and benchmark coverage moved 14/193 → 14/193. Zero. Forty-five correct new entries bought no measured improvement, precisely because the guard forbids harvesting anything a benchmark item would match. That is the design working, not a limitation to fix.
The same rule governs what we take from others: we map other benchmarks' coverage, never their items. Nothing in this estate has been ingested from another benchmark's dataset, every source's licence is recorded asVERIFIED_PERMISSIVE /VERIFIED_RESTRICTED /UNVERIFIED, and the default is UNVERIFIED — which is not a synonym for "probably fine". An audit refuses to pass any source marked ingestible without a verified licence.
What this is not
This governs provenance, not correctness. The pipeline has shipped a wrong legal answer carrying a valid Article 50 marking and a clean signed receipt. An attested answer is attested, never verified.
UNCERTIFIED is the default. No competent authority exists to confer EU AI Act conformity, so neither can we. Nothing here is a certification, and we hold no accreditation.
Models are subjects here, not components. We do not host, merge or average other people's models — we score them. Scoring a public checkpoint needs no permission and is free; scoring a company's deployed system needs a contract, because it sits behind their auth. Anyone claiming to have absorbed the field is describing something that would produce nothing new.
Every result is valid only for the item set it was measured on. Item-set fingerprints are published so you can tell when a score has gone stale. Ours had.
Tokens per correct verdict
The production number nobody else publishes. Every token does triple duty: benchmark (token efficiency on governance work), evidence (daily accumulated compliance behaviour), fuel (training pairs + KB rows, practice split only). The flywheel never trains on what it scores — the held-out split is enforced by a salted content hash and export_fuel() raises if a held-out item ever reaches it. Without this, the flywheel eats itself — that's the Leaderboard Illusion, and our own defbench already proved the local version.
| model | practice acc | held-out acc | overfit gap | guard | exported | when |
|---|---|---|---|---|---|---|
| qwen2.5:1.5b | 0.550 (n=780) | — (n=0) | 0.000 | export 780 pairs | 780p / 780kb | 2026-08-15T16:07:05.674460+00:00 |
| qwen2.5:1.5b | 0.550 (n=720) | — (n=0) | 0.000 | export 720 pairs | 720p / 720kb | 2026-08-15T16:07:05.714743+00:00 |
| qwen2.5:1.5b | 0.552 (n=703) | — (n=0) | 0.000 | export 703 pairs | 703p / 703kb | 2026-08-15T16:07:05.725046+00:00 |
| qwen2.5:1.5b | 0.550 (n=680) | — (n=0) | 0.000 | export 680 pairs | 680p / 680kb | 2026-08-15T16:07:05.739844+00:00 |
| qwen2.5:1.5b | 0.484 (n=580) | — (n=0) | 0.000 | export 580 pairs | 580p / 580kb | 2026-08-15T16:07:05.872991+00:00 |
| qwen2.5:1.5b | 0.450 (n=100) | — (n=0) | 0.000 | export 100 pairs | 100p / 100kb | 2026-08-15T16:07:05.883180+00:00 |