measured · Council-34:latest

GovBench

governance · seat of the instrument: Brussels · EU AI Act (Reg. 2024/1689)

EU AI Act risk-tier classification. Graded deterministically — a regex extracts the label and the result is scored by macro-F1. No model judges another model.

What has been measured

modelmacro-F1accuracy 95% CIunreadablenharness
Council-34:latest0.3810.515 [0.451, 0.578]0%237measure_full.py
falcon3:7b0.3530.426 [0.365, 0.490]0%237measure_full.py
qwen2.5:1.5b0.2910.430 [0.369, 0.494]0%237measure_full.py
falcon3:7b0.468n<30 — not quotable0%24measure_robust2.py
Council-34:latest0.345n<30 — not quotable25%24measure_robust2.py
qwen2.5:1.5b0.343n<30 — not quotable0%24measure_robust2.py
Council-34:latest0.386n<30 — not quotable4%measure.py
This axis is quotable. n has reached usable_n = 30, so a confidence interval is published. Overlapping intervals mean the models are not separated — a lead is not a separation.
3 runs dropped — Council-34:latest, falcon3:7b, qwen2.5:1.5b at 97–100% instrument error (all URLError). A dropped connection is not a wrong answer, so these contribute no score.

The chain — 8 of 10 complete

A measure with a dataset and nothing else is one link, not a chain. This list is measured live, not asserted.

Open the Space ↗The dataset ↗Kaggle ↗

Run it yourself

Run GovBench yourself

The same 24 items 7 models answered, graded by the same deterministic rule. The models measured on these items are ranked with you when you finish. No sign-up, nothing leaves your browser.

Items: csoai/gspc-gov · grading is a regex label read plus macro-F1, identical to the published harness · measurement, not certification, and not legal advice.