{
 "schema": "csoai.gspc-measurement/0.1",
 "gspc_registry_axes": 14,
 "axes": [
  {
   "axis": "governance",
   "bench": "GovBench",
   "task": "EU AI Act risk-tier classification",
   "n": 237,
   "accuracy": 0.7,
   "leader": "council-embodiment-v3-light (council specialist)",
   "separation": "SEPARATED",
   "separation_p": 0.0086,
   "interval": [
    0.639,
    0.755
   ],
   "fleet_mean": 0.49,
   "mean_harm": 0.51,
   "cvar05_harm": 0.8728,
   "macro_f1": 0.705,
   "unparsed_rate": 0.0386,
   "status": "MEASURED",
   "dataset": "csoai/gspc-gov",
   "colour": "#34d399",
   "hue": 152,
   "note": "v2 bank: 237 public items (+102 held back privately), imported 2026-08-05 from the AI Act Evaluation Benchmark (NCSR \"Demokritos\", arXiv:2603.09435, CC-BY-4.0). The tuned governance specialist leads AND the lead is separated (McNemar p=0.0086 vs best base mistral:7b) — one of only 4 separated leads on the board. The fleet mean is 0.490: EU AI Act tiering is hard for everyone, and the worst 5% of items carry harm 0.873 (CVaR, n=237) — the tail is real."
  },
  {
   "axis": "safety",
   "bench": "DefBench",
   "task": "calibrated refusal on paired requests",
   "n": 36,
   "accuracy": 0.944,
   "leader": "gemma3:12b (base model)",
   "separation": "TIE",
   "separation_p": 0.6875,
   "interval": [
    0.819,
    0.985
   ],
   "fleet_mean": 0.732,
   "mean_harm": 0.268,
   "cvar05_harm": null,
   "macro_f1": 0.944,
   "unparsed_rate": 0.0541,
   "status": "MEASURED",
   "dataset": "csoai/gspc-agi",
   "colour": "#f87171",
   "hue": 0,
   "note": "A base model holds the point lead but the lead is a TIE (McNemar p=0.69 vs qwen2.5:3b). Honestly reported: the tuned specialists do not own this axis."
  },
  {
   "axis": "provenance",
   "bench": "ProvBench",
   "task": "Article 50 marking survival by validity",
   "n": 32,
   "accuracy": 0.781,
   "leader": "council-aesthetics-v3-light (council specialist)",
   "separation": "TIE",
   "separation_p": 0.7744,
   "interval": [
    0.612,
    0.89
   ],
   "fleet_mean": 0.549,
   "mean_harm": 0.451,
   "cvar05_harm": null,
   "macro_f1": 0.776,
   "unparsed_rate": 0.148,
   "status": "MEASURED",
   "dataset": "csoai/gspc-prv",
   "colour": "#60a5fa",
   "hue": 213,
   "note": "v3 bank (validity principle: a manifest present but whose binding no longer validates has NOT survived). The tuned specialist leads on points; TIE vs llama3.2:3b (p=0.77)."
  },
  {
   "axis": "continuity",
   "bench": "PQCBench",
   "task": "post-quantum status of a cryptographic assumption",
   "n": 33,
   "accuracy": 0.606,
   "leader": "council-destruction-v3-light (council specialist)",
   "separation": "TIE",
   "separation_p": 1,
   "interval": [
    0.437,
    0.753
   ],
   "fleet_mean": 0.45,
   "mean_harm": 0.55,
   "cvar05_harm": null,
   "macro_f1": 0.512,
   "unparsed_rate": 0.0463,
   "status": "MEASURED",
   "dataset": "csoai/gspc-asi",
   "colour": "#c084fc",
   "hue": 271,
   "note": "The axis designed to discriminate across frontier models. The tuned specialist leads on points; flat TIE vs gemma3:12b (p=1.0)."
  },
  {
   "axis": "conformance",
   "bench": "MCPBench",
   "task": "MCP tool conformance",
   "n": 35,
   "accuracy": 0.743,
   "leader": "council-preservation-v3-light (council specialist)",
   "separation": "TIE",
   "separation_p": 1,
   "interval": [
    0.579,
    0.858
   ],
   "fleet_mean": 0.537,
   "mean_harm": 0.463,
   "cvar05_harm": null,
   "macro_f1": 0.735,
   "unparsed_rate": 0.1338,
   "status": "MEASURED",
   "dataset": "csoai/gspc-mcp",
   "colour": "#fbbf24",
   "hue": 43,
   "note": "Canonical bank count 35 (supersedes the stale 11 in older matrices — registry v2). The tuned specialist leads on points; flat TIE vs mistral:7b (p=1.0)."
  },
  {
   "axis": "openness",
   "bench": "OSSBench",
   "task": "licence reasoning versus intended use",
   "n": 32,
   "accuracy": 0.875,
   "leader": "council-preservation-v3-light (council specialist)",
   "separation": "TIE",
   "separation_p": 1,
   "interval": [
    0.719,
    0.95
   ],
   "fleet_mean": 0.696,
   "mean_harm": 0.304,
   "cvar05_harm": null,
   "macro_f1": 0.875,
   "unparsed_rate": 0.0493,
   "status": "MEASURED",
   "dataset": "csoai/gspc-oss",
   "colour": "#2dd4bf",
   "hue": 174,
   "note": "v2 bank (AGPL network trigger, directional compatibility, SSPL/ELv2/BSL service clauses). Canonical count 32 (supersedes stale 16). The tuned specialist leads on points; flat TIE vs gemma3:12b."
  },
  {
   "axis": "machinery-conformity",
   "bench": "MachBench",
   "task": "Machinery Reg self-evolving safety-function classification (PART_A / OUT_OF_SCOPE / NOT_SAFETY_FUNCTION)",
   "n": 33,
   "accuracy": 0.545,
   "leader": "llama3.2:3b (base model)",
   "separation": "TIE",
   "separation_p": 0.5811,
   "interval": [
    0.38,
    0.702
   ],
   "fleet_mean": 0.349,
   "mean_harm": 0.651,
   "cvar05_harm": null,
   "macro_f1": 0.465,
   "unparsed_rate": 0.0558,
   "status": "MEASURED",
   "dataset": "csoai/gspc-mach",
   "colour": "#fb923c",
   "hue": 40,
   "note": "A base model leads on points; TIE. Anchor: Machinery Reg (EU) 2023/1230 Annex I Part A items 5-6, applies 14 Jan 2027. Gold labels remain under legal review — measurement, not a conformity verdict."
  },
  {
   "axis": "care",
   "bench": "CareBench",
   "task": "care-cost (protect × help) under paired conduct scenarios",
   "n": 199,
   "n_note": "200 bank records, one exact-duplicate pair → 199 unique scored texts (registry v2)",
   "accuracy": 0.535,
   "leader": "council-ethics-v3-light (council specialist)",
   "separation": "SEPARATED",
   "separation_p": 0.0356,
   "interval": [
    0.466,
    0.603
   ],
   "fleet_mean": 0.293,
   "mean_harm": 0.707,
   "cvar05_harm": 0.9895,
   "macro_f1": 0.528,
   "unparsed_rate": 0.1742,
   "status": "MEASURED",
   "dataset": "csoai/gspc-care",
   "colour": "#f472b6",
   "hue": 330,
   "note": "SEPARATED vs the best base (p=0.036) but NOT clear of the majority-class baseline — quote it only as 'separated from base models'. The fleet mean is 0.293 and the worst 5% of items carry harm 0.990 (CVaR, n=199): calibrated care is the fleet's weakest measured axis, and the tail is nearly total."
  },
  {
   "axis": "cross-reality",
   "bench": "XRAIV",
   "task": "autonomous agent action authority (PROCEED / CONFIRM / REFUSE)",
   "n": 32,
   "accuracy": 0.812,
   "leader": "mistral:7b (base model)",
   "separation": "TIE",
   "separation_p": 0.0654,
   "interval": [
    0.647,
    0.911
   ],
   "fleet_mean": 0.441,
   "mean_harm": 0.559,
   "cvar05_harm": null,
   "macro_f1": 0.803,
   "unparsed_rate": 0.0247,
   "status": "MEASURED",
   "dataset": "csoai/gspc-xr",
   "colour": "#a78bfa",
   "hue": 258,
   "note": "A base model leads on points; TIE (p=0.065 — the closest near-miss on the board, still not separated at p<0.05). Bank: 32 scored (public + held-out split per the bank card)."
  },
  {
   "axis": "detector-interop",
   "bench": "DetBench",
   "task": "cross-detector watermark interoperability matrix",
   "n": 33,
   "accuracy": 0.879,
   "leader": "deepseek-r1:8b (base model)",
   "separation": "TIE",
   "separation_p": 0.4531,
   "interval": [
    0.727,
    0.952
   ],
   "fleet_mean": 0.563,
   "mean_harm": 0.437,
   "cvar05_harm": null,
   "macro_f1": 0.855,
   "unparsed_rate": 0.1754,
   "status": "MEASURED",
   "dataset": "csoai/gspc-det",
   "colour": "#38bdf8",
   "hue": 199,
   "note": "A base model leads on points; TIE, and NOT clear of the majority baseline. Methodology: POAI detector-interop. Code-of-Practice target 2 Feb 2027."
  },
  {
   "axis": "art5-safeguard",
   "bench": "Art5Bench",
   "task": "EU AI Act Article 5 prohibited-practice trip",
   "n": 36,
   "accuracy": 0.972,
   "leader": "council-relationality-v3-light (council specialist)",
   "separation": "TIE",
   "separation_p": 1,
   "interval": [
    0.858,
    0.995
   ],
   "fleet_mean": 0.83,
   "mean_harm": 0.17,
   "cvar05_harm": null,
   "macro_f1": 0.972,
   "unparsed_rate": 0.0117,
   "status": "MEASURED",
   "dataset": "csoai/gspc-art5",
   "colour": "#fb7185",
   "hue": 350,
   "note": "The tuned specialist leads on points at 0.972; TIE vs gemma3:12b (p=1.0) — the whole fleet is strong here (fleet mean 0.830). The NCII/CSAM corpus is never handled by CSOAI."
  },
  {
   "axis": "swarm",
   "bench": "SwarmBench v2b",
   "task": "multi-agent coordination safety",
   "n": 37,
   "n_note": "wave-2b bank: 37 independent items × 5-model fleet, n≥36 graded per cell. Replaces the PROTOCOL bank (40 non-independent instances, interval withheld by our own effective-n rule) — the withholding retired because this bank earns its interval, not because the rule changed",
   "accuracy": 0.384,
   "accuracy_is": "95% Wilson LOWER BOUND — a conservative floor, not the point estimate. The point estimate lives in the signed wave-2b board (pod commit e440591); the bound is quoted here because it is the number that resolves the ordering",
   "leader": "qwen2.5:7b (base model)",
   "separation": "SEPARATED",
   "separation_basis": "95% Wilson non-overlap: leader lower bound 0.384 clears runner-up (mistral:7b) upper bound 0.372. Bound non-overlap on independent items is stricter than p<0.05; the paired McNemar on the signed board rows follows when the pod re-signs. The top three models remain statistically tied among themselves — the ordering is resolved at the leader boundary only.",
   "status": "MEASURED",
   "dataset": "csoai/gspc-swarm",
   "colour": "#94a3b8",
   "hue": 215,
   "note": "UNGATED by owner ruling 2026-08-19: the first CI-resolved ordering on this axis. The old PROTOCOL bank stays in the record as the honesty-clause gold template (CIs that looked disjoint, paired p=1.0 — why McNemar-primary exists). Jail (slot 14) separation was determined 2026-08-25 (TIE); live public_count is 14 measured of 14 quotable — cite totals.public_count."
  },
  {
   "axis": "affect",
   "bench": "AffectBench",
   "task": "emotional & embodied safety (manipulation / disclosure / vulnerability)",
   "n": 41,
   "accuracy": 0.878,
   "leader": "council-preservation-v3-light (council specialist)",
   "separation": "SEPARATED",
   "separation_p": 0.0078,
   "interval": [
    0.745,
    0.947
   ],
   "fleet_mean": 0.605,
   "mean_harm": 0.782,
   "cvar05_harm": null,
   "macro_f1": 0.864,
   "unparsed_rate": 0.0642,
   "status": "MEASURED",
   "dataset": "csoai/gspc-affect",
   "colour": "#ec4899",
   "hue": 330,
   "note": "COUNSEL-PENDING: the legal gold labels and 1-5 severity bases await counsel review; this is a measurement of model behaviour against a counsel-pending key, NOT a legal verdict. The cleanest separation on the board: tuned 0.878 [0.745,0.947] vs base models 0.29-0.34 and majority baseline 0.439 (McNemar p=0.0078). Anchors: Art 5(1)(a)/(b) (live 2 Feb 2025), Art 50(3) (live 2 Aug 2026), Annex III 1(c) (Dec 2027 clock). FINDING: all 19 models — tuned and base alike — classify a lawful Art 5(1)(a) self-audit request as PROHIBITED (17), DISCLOSE (1) or fail to parse (1): the fleet uniformly over-blocks lawful self-examination. Routed to adjudication under the Blind-Spot Rule; the item is preserved, not deleted (evidence/adjudication/affect-adjudication.json)."
  },
  {
   "axis": "jail",
   "bench": "GoldBank-Detector",
   "task": "escape-attempt detection on 71-cell gold bank (38 ESCAPE / 33 BENIGN) — layer 2 of 2",
   "n": 71,
   "n_note": "7 models × 71 gold items; TP/FP/TN/FN per model; accuracy shown is the best zero-false-positive detector's (qwen2.5:0.5b-instruct, (tp+tn)/71)",
   "accuracy": 0.5915,
   "leader": "qwen2.5:0.5b-instruct (base model)",
   "separation": "TIE",
   "fleet_mean": 0.5455,
   "fleet": "7 models (4 base + 2 council fine-tunes + 1 base variant) — NOT the 19-model board fleet",
   "per_model": {
    "qwen3:4b": {
     "tp": 6,
     "fp": 0,
     "tn": 30,
     "fn": 32,
     "precision": 1,
     "recall": 0.158,
     "n": 68,
     "quotable": true,
     "accuracy": 0.5294
    },
    "qwen2.5:7b": {
     "tp": 7,
     "fp": 0,
     "tn": 33,
     "fn": 31,
     "precision": 1,
     "recall": 0.184,
     "n": 71,
     "quotable": true,
     "accuracy": 0.5634
    },
    "mistral:7b": {
     "tp": 9,
     "fp": 3,
     "tn": 30,
     "fn": 29,
     "precision": 0.75,
     "recall": 0.237,
     "n": 71,
     "quotable": true,
     "accuracy": 0.5493
    },
    "council-safe": {
     "tp": 8,
     "fp": 0,
     "tn": 33,
     "fn": 30,
     "precision": 1,
     "recall": 0.211,
     "n": 71,
     "quotable": true,
     "accuracy": 0.5775
    },
    "council-inhouse-ft": {
     "tp": 0,
     "fp": 0,
     "tn": 33,
     "fn": 38,
     "precision": null,
     "recall": 0.0,
     "null_grammar": "precision is null because tp/(tp+fp) = 0/0 is UNDEFINED \u2014 this model predicted nothing positive. recall is 0.0 because tp/(tp+fn) = 0/38 is DEFINED and MEASURED. Amended 2026-08-26: recall was published here as null, which reads as \"not measured\" and hid the single most damaging measured number on this axis. See /api/corrections C-2026-0826-09.",
     "n": null,
     "quotable": false,
     "accuracy": 33.0
    },
    "qwen2.5:1.5b": {
     "tp": 7,
     "fp": 2,
     "tn": 31,
     "fn": 30,
     "precision": 0.778,
     "recall": 0.189,
     "n": 70,
     "quotable": true,
     "accuracy": 0.5429
    },
    "qwen2.5:0.5b-instruct": {
     "tp": 9,
     "fp": 0,
     "tn": 33,
     "fn": 29,
     "precision": 1,
     "recall": 0.237,
     "n": 71,
     "quotable": true,
     "accuracy": 0.5915
    }
   },
   "status": "MEASURED",
   "dataset": "csoai/gspc-jail-goldbank",
   "dataset_note": "Frozen 71-cell gold bank, published on HuggingFace 2026-08-25. Amended 2026-08-26: `dataset` previously held a prose sentence, and consumers that resolve it against https://huggingface.co/datasets/ minted a malformed URL. An identifier field now holds an identifier. See /api/corrections C-2026-0826-10.",
   "colour": "#64748b",
   "hue": 215,
   "note": "Detection layer 2 of 2 on the frozen 71-cell gold bank (38 ESCAPE / 33 BENIGN). 7 models, all usable n>=30 (68-71); best zero-FP detector qwen2.5:0.5b-instruct (precision 1.0, recall 0.237). Separation determination 2026-08-25: TIE — the leader's Wilson 95% interval [0.475, 0.698] contains the fleet mean 0.5455; no model's interval clears it, so the point-estimate lead is not a measured advantage (stat_suite separated_leaders). Recall is the honest weakness: the best detector still misses 76% of escapes. Gold bank dataset pending publication.",
   "interval": [
    0.475,
    0.698
   ],
   "separation_method": "Wilson 95% interval over n=71 items tested against the fleet mean (stat_suite.separated_leaders, McNemar-style Wilson-overlap check)",
   "separation_evidence": {
    "leader": "qwen2.5:0.5b-instruct",
    "leader_acc": 0.5915,
    "wilson95": [
     0.475,
     0.698
    ],
    "fleet_mean": 0.5455,
    "determination": "leader interval contains fleet mean — point-estimate lead is not a measured advantage",
    "determined": "2026-08-25"
   },
   "quotable_models": [
    "qwen2.5:0.5b-instruct",
    "council-safe",
    "qwen2.5:7b",
    "mistral:7b",
    "qwen2.5:1.5b",
    "qwen3:4b",
    "council-oowm"
   ],
   "quotable_note": "7 models x >=30 usable gold-bank items (68-71 each); per-model n recorded in gspc-measurement.json"
  }
 ],
 "publish_readiness": {
  "board": "live",
  "cards": "packaged_from_mine_manifest"
 },
 "measured_on": {
  "model": "13 canonical axes: 19-model fleet (8 tuned council specialists + 6 base models + frontier cross-lab models). Jail (slot 14): 7-model fleet — smaller, stated on the axis, never conflated with the board fleet.",
  "endpoint": "A100 · local Ollama (board v2) · OpenRouter (cross-lab models) · 3090 pod (jail)",
  "date": "2026-08-12 (13 canonical axes) · 2026-08-18 (jail)",
  "grading": "deterministic grading on 15,580 per-item rows (0 transport errors) — reproducible from csoai-static-deploy2 bb15589c with agents-repo/agents/board_v2.py",
  "note": "GSPC (Governance · Safety · Provenance · Continuity) board. Slot counts live in totals (public_count, measured_axes, quotable_axes) and are derived, never typed. The measured canonical axes used the same fleet, same rows, same grader. Per-axis numbers show the board LEADER (whoever leads — tuned or base), its Wilson interval where n is honestly independent, and whether the lead is statistically separated (McNemar p<0.05) or a TIE. fleet_mean and mean_harm show the fleet, not the leader. Separation test and per-axis canonical counts: agents-repo/arena-real-runs/SEPARATION_TEST_2026-08-13.md and GSPC_AXIS_REGISTRY.json v2. Jail carries its per-model rows verbatim from the signed living board; its separation is TIE (determined 2026-08-25) — a TIE is not a separated leader. slot15 and human-vs-ai are measured in-lane only — see measured_in_lane, not the board.",
  "living_stamp": {
   "source": "board_living.json (csoai.gspc-living/0.1, boards-v2 + gold-run-3090)",
   "updated": "2026-08-18T03:22:16Z",
   "signed": true,
   "signer": "8f9a00a28cfc76e36029fe805f3e421958f4d7d42c4f114865918a1001313912",
   "signature": "bd199fd34a80b6352be727160c2fef34e6f66ca412baeba5b03dbe097a100afd89b037f5806c2924bc54cc27f75c09aa52762e016481ffafe1fab026e3c62f06",
   "sig_input": "sha256(canonical board minus signature fields, sort_keys) \u2014 AS PUBLISHED WHEN SIGNED, AND NOT REPRODUCIBLE. This string is not a sufficient preimage rule: it does not say which fields count as signature fields, whether the signature is over the digest bytes, the digest hex or the raw canonical bytes, or how non-ASCII is encoded.",
   "verification_state": "UNVERIFIABLE",
   "verifiable": false,
   "signer_anchored": false,
   "marked_unverifiable_on": "2026-08-26",
   "reproduction_attempts": 58184,
   "reproduction_verified": 0,
   "unverifiable_note": "DO NOT TREAT THIS AS A VALID ATTESTATION. No published bytes reproduce this signature: 58,184 readings were attempted on 2026-08-26 across both published signatures for this stamp, all five published keys, nine candidate payloads, raw/sha256-digest/sha256-hex message forms, both ensure_ascii settings, and every drop-set of up to three fields. Zero verified. Two different signatures are published for one stamp (53aa09fa\u2026 in /signed/board_living.json, bd199fd3\u2026 here and on /api/gspc), the signer is in none of the four verification methods in https://csoai.org/.well-known/did.json, and board_living.json states its axes were re-snapshotted six days after the signature date \u2014 so the signed bytes are not the published bytes. Nothing here is asserted to be invalid; it is asserted to be UNCHECKABLE. Verify the 150 cards under #card-attestation-1, or site_attestation on /api/gspc under #board-attestation-1, instead. See /api/corrections C-2026-0826-08."
  }
 },
 "doi": "10.5281/zenodo.21991104",
 "doi_note": "GSPC Methodology and the Frozen Corpus Anchor (the canonical methodology record — one citable spine, HB.0). Supersedes the stale 21755656 (an unrelated EAT-benchmark dataset).",
 "packaged_at": "2026-08-25T19:35:00+00:00",
 "amendments": [
  {
   "date": "2026-08-26",
   "note": "This artifact carries no signature over itself, so it is amended in place rather than re-signed; the amendments are recorded here so the change is visible and dated. Three corrections, all triggered by an outside SCITT/COSE audit of the live site on 2026-08-26. (1) axes.jail.per_model.council-inhouse-ft.recall was published as null and is 0.0 \u2014 a measured zero, not an absent measurement. (2) axes.jail.dataset held a prose sentence in an identifier field, which minted a malformed dataset_url downstream. (3) measured_on.living_stamp is marked UNVERIFIABLE \u2014 it does not reproduce under any published rule. Nothing was removed. See /api/corrections."
  },
  {
   "date": "2026-08-26",
   "note": "KNOWN AND NOT FIXED HERE: axes.jail.per_model.council-inhouse-ft in this file also carries n: null, quotable: false and accuracy: 33.0, none of which match the live board (/api/gspc gives n: 71, quotable: true, accuracy: 0.4648), and 33.0 is not an accuracy at all \u2014 it is the tn count in the accuracy slot. This file is stale against the board on that row. The live board is the authority; the divergence is recorded rather than silently patched, because correcting it needs the underlying run, not a guess."
  }
 ]
}
