GSPC board v2 — 13 measurement axes
Measured 2026-08-12 · 19-model fleet · 15,580 per-item rows · schema csoai.gspc-axes/0.3 · DOI 10.5281/zenodo.21755656. A leader is the highest point estimate; separation is McNemar p<0.05 on discordant items vs the best base model. 3 of 13 separated, 10 ties. Ties are honest ties — a point-estimate lead is not a measured win.
Reconciliation notice (2026-08-16). A corrected honest register records a base model (mistral:7b) leading the council-specialist on the governance axis in a later re-measurement. This board stands as the documented 2026-08-12 sweep until the owner-lane reconciliation decides whether the API leader table is corrected or annotated. A measurement body publishes disagreements, never hides them.
| Axis | Leader | Accuracy (95% CI) | n | Separation |
|---|---|---|---|---|
| governance | council specialist:governance-v3 | 70.0% [63.9%, 75.5%] | 237 | SEPARATED · p=0.0086 |
| affect | council specialist:preservation-v3 | 87.8% [74.5%, 94.7%] | 41 | SEPARATED · p=0.0078 |
| care | council specialist:ethics-v3 | 53.5% [46.6%, 60.3%] | 199 | SEPARATED · p=0.0356 |
| art5-safeguard | council specialist:relationality-v3 | 97.2% [85.8%, 99.5%] | 36 | TIE · p=1 |
| safety | gemma3:12b (base model) | 94.4% [81.9%, 98.5%] | 36 | TIE · p=0.6875 |
| detector-interop | deepseek-r1:8b (base model) | 87.9% [72.7%, 95.2%] | 33 | TIE · p=0.4531 |
| openness | council specialist:preservation-v3 | 87.5% [71.9%, 95.0%] | 32 | TIE · p=1 |
| cross-reality | mistral:7b (base model) | 81.2% [64.7%, 91.1%] | 32 | TIE · p=0.0654 |
| provenance | council specialist:aesthetics-v3 | 78.1% [61.2%, 89.0%] | 32 | TIE · p=0.7744 |
| conformance | council specialist:preservation-v3 | 74.3% [57.9%, 85.8%] | 35 | TIE · p=1 |
| continuity | council specialist:destruction-v3 | 60.6% [43.7%, 75.3%] | 33 | TIE · p=1 |
| machinery-conformity | llama3.2:3b (base model) | 54.5% [38.0%, 70.2%] | 33 | TIE · p=0.5811 |
| swarm | qwen2.5:0.5b-instruct (base model) | 97.5% (no interval — effective-n rule) | 40 | TIE · p=1 |
819 items across 13 axes. swarm withholds its interval by the effective-n rule (3 unique prompts, 40 non-independent instances). Recompute the board live at councilof.ai/api/gspc. Measurement, not certification.
click to launch global
Regulator Lens
Cross-framework compliance · audit trail
AIR-Bench refusal sweep
A separate, secondary measurement from a one-off AIR-Bench sweep (2026-08-04) — not board v2, not refreshed. Refusal rate is over measured prompts only; UNMEASURED ≠ fail.
Static snapshot of the one-off sweep measured 2026-08-04 — not refreshed. The growth figure is illustrative, not a measurement.