measured · Council-34:latest
governance · seat of the instrument: Brussels · EU AI Act (Reg. 2024/1689)
EU AI Act risk-tier classification. Graded deterministically — a regex extracts the label and the result is scored by macro-F1. No model judges another model.
| model | macro-F1 | accuracy 95% CI | unreadable | n | harness |
|---|---|---|---|---|---|
| Council-34:latest | 0.381 | 0.515 [0.451, 0.578] | 0% | 237 | measure_full.py |
| falcon3:7b | 0.353 | 0.426 [0.365, 0.490] | 0% | 237 | measure_full.py |
| qwen2.5:1.5b | 0.291 | 0.430 [0.369, 0.494] | 0% | 237 | measure_full.py |
| falcon3:7b | 0.468 | n<30 — not quotable | 0% | 24 | measure_robust2.py |
| Council-34:latest | 0.345 | n<30 — not quotable | 25% | 24 | measure_robust2.py |
| qwen2.5:1.5b | 0.343 | n<30 — not quotable | 0% | 24 | measure_robust2.py |
| Council-34:latest | 0.386 | n<30 — not quotable | 4% | — | measure.py |
A measure with a dataset and nothing else is one link, not a chain. This list is measured live, not asserted.
The same 24 items 7 models answered, graded by the same deterministic rule. The models measured on these items are ranked with you when you finish. No sign-up, nothing leaves your browser.
Items: csoai/gspc-gov · grading is a regex label read plus macro-F1, identical to the published harness · measurement, not certification, and not legal advice.