Governance lens

Does it comply with statute?

runs in your browser · nothing is sent anywhere

Describe an AI system. The instrument reads it against Article 5, Article 6 and Annex III and tells you which provisions bind — deterministically, with the provision cited, offline.

Battles — why they are coming, and why they are not here yet

not built

0 of 15 governance dimensions currently resolve. Every model is statistically tied because the minimum detectable effect is ≈63 points against observed margins of 1–15. More absolute-scored items barely move that: going from 6 to 11 items per dimension moved interval resolution by exactly zero.

Pairwise comparison is more statistically efficient than absolute scoring, which is why arena-style ranking resolves where ours does not. But the mechanism that gets them there is one we cannot use: Arena-Rank (Apache-2.0) is a Bradley-Terry model over human preference votes — crowd judgement, and our first design law is that every primary score is deterministic with no judge. So we borrow the presentation — a score with an explicit 95% CI and an explicit n — and reject the Elo. Any battles here would resolve a pairwise verdict from the five deterministic predicates against a hashed provision, never from a vote.

The line battles must not cross

Public votes may never enter the benchmark. GovBench items are exact-match scored, so harvesting site traffic back into them is circular and would void every published score. Battle votes will be a separate, disclosed, opt-in signal with their own leaderboard — never merged into GovBench, and never used as training data.