Measurement Leaderboard

Static snapshot · one-off sweep measured 2026-08-04 · refusal rate over measured prompts only · UNMEASURED ≠ fail

GSPC board v2 — 13 measurement axes

Measured 2026-08-12 · 19-model fleet · 15,580 per-item rows · schema csoai.gspc-axes/0.3 · DOI 10.5281/zenodo.21755656. A leader is the highest point estimate; separation is McNemar p<0.05 on discordant items vs the best base model. 3 of 13 separated, 10 ties. Ties are honest ties — a point-estimate lead is not a measured win.

Reconciliation notice (2026-08-16). A corrected honest register records a base model (mistral:7b) leading the council-specialist on the governance axis in a later re-measurement. This board stands as the documented 2026-08-12 sweep until the owner-lane reconciliation decides whether the API leader table is corrected or annotated. A measurement body publishes disagreements, never hides them.

AxisLeaderAccuracy (95% CI)nSeparation
governancecouncil specialist:governance-v370.0% [63.9%, 75.5%]237
SEPARATED · p=0.0086
affectcouncil specialist:preservation-v387.8% [74.5%, 94.7%]41
SEPARATED · p=0.0078
carecouncil specialist:ethics-v353.5% [46.6%, 60.3%]199
SEPARATED · p=0.0356
art5-safeguardcouncil specialist:relationality-v397.2% [85.8%, 99.5%]36
TIE · p=1
safetygemma3:12b (base model)94.4% [81.9%, 98.5%]36
TIE · p=0.6875
detector-interopdeepseek-r1:8b (base model)87.9% [72.7%, 95.2%]33
TIE · p=0.4531
opennesscouncil specialist:preservation-v387.5% [71.9%, 95.0%]32
TIE · p=1
cross-realitymistral:7b (base model)81.2% [64.7%, 91.1%]32
TIE · p=0.0654
provenancecouncil specialist:aesthetics-v378.1% [61.2%, 89.0%]32
TIE · p=0.7744
conformancecouncil specialist:preservation-v374.3% [57.9%, 85.8%]35
TIE · p=1
continuitycouncil specialist:destruction-v360.6% [43.7%, 75.3%]33
TIE · p=1
machinery-conformityllama3.2:3b (base model)54.5% [38.0%, 70.2%]33
TIE · p=0.5811
swarmqwen2.5:0.5b-instruct (base model)97.5% (no interval — effective-n rule)40
TIE · p=1

819 items across 13 axes. swarm withholds its interval by the effective-n rule (3 unique prompts, 40 non-independent instances). Recompute the board live at councilof.ai/api/gspc. Measurement, not certification.

click to launch global

Regulator Lens

Regulator Lens

Cross-framework compliance · audit trail

never auto-resolved

AIR-Bench refusal sweep

A separate, secondary measurement from a one-off AIR-Bench sweep (2026-08-04) — not board v2, not refreshed. Refusal rate is over measured prompts only; UNMEASURED ≠ fail.

1,398
Prompts Measured
8
Subjects
4
Instruments
one-off
Snapshot Cadence
EU AI Act
Largest Instrument
+24%
Monthly Growth

Static snapshot of the one-off sweep measured 2026-08-04 — not refreshed. The growth figure is illustrative, not a measurement.

gpt-oss-20b
Measured
136 cases73.9% accuracy
184
points
gpt-oss-120b
Measured
226 cases66.9% accuracy
338
points
gemma-4-26b-a4b-it
Measured
119 cases51.5% accuracy
231
points
#4
qwen3.6-27b
Measured
96 cases38.4% accuracy
250
points
#5
llama-3.1-8b-instant
Measured
39 cases30.7% accuracy
127
points
#6
llama-3.3-70b-versatile
Measured
46 cases28.9% accuracy
159
points
#7
allam-2-7b
Measured
25 cases23.6% accuracy
106
points
#8
gemma-4-31b-it
LOW_N
1 cases33.3% accuracy
3
points

Achievements

First 1,000 Measured
Sweep crossed 1,000 measured prompts
1 holders
8 Subjects
Eight model families under measurement
8 holders
4 Instruments Live
127 provisions under continuous hash watch
4 holders
Kaggle Flag
csoai-corpus-baselines public on Kaggle
1 holders
Ledger Published
Signed measurement ledger from the 2026-08-04 sweep on HF
1 holders

How to Participate

1

Submit AI safety incident reports through our platform or browser extension

2

Earn points when your reports are verified by certified analysts

3

Climb the leaderboard and unlock achievements

4

Become a certified analyst to review and verify reports