Measurement, never certification
AILuminate for chat. GSPC beyond chat.
AILuminate (MLCommons AI Safety v1.1, 11 Feb 2025, arXiv 2503.05731) is the industry's chat-hazard benchmark, and a good one. GSPC is a 22-axis measurement board. They are not rivals — one is an attachment to the other. This page says exactly where each is strong, and exactly how results from the first may ride on the second without ever pretending to be it.
| Dimension | AILuminate v1.1 | GSPC |
|---|---|---|
| What it measures | Chat-style hazard response — one conversational surface | 22 axes across two families: 14 behavioural (governance, safety, provenance, continuity, conformance, openness, machinery, care, cross-reality, detector-interop, art5-safeguard, swarm, affect, jail) + 8 financial/domain |
| What a result is | A 5-tier grade (Poor … Excellent) per hazard category | A measurement: n, accuracy, Wilson interval, a leader, a separation verdict — or an honestly empty UNMEASURED cell |
| Reference baseline | Relative to a floating reference model — a grade is not comparable across time | Frozen banks, pinned datasets, replayable runs — a number today can be checked against itself tomorrow |
| Prompt sets | Practice set + confidential official set with rotating reserve prompts (the right anti-Goodhart shape) | Published frozen banks + held-out practice discipline; the same doctrine, applied beyond chat |
| Reach | 12 hazard categories (14 codes) of chat risk | Behavioural + financial + regulatory + provenance surfaces a chat benchmark does not reach — reserve attestation, custody disclosure, distribution integrity among them |
| Output you can verify | A grade reported by the framework | A signed card, content-addressed, included in a public Merkle root — verification is free forever |
Where AILuminate is deeper — said out loud
- Chat-hazard granularity: 12 categories with per-category practice prompts are deeper on the chat surface than any single GSPC axis tries to be.
- Set governance: the practice/official split with a monitored reserve is the anti-Goodhart template the estate's own holdout doctrine copies.
- Institutional weight: MLCommons membership gives AILuminate adoption a chat benchmark earns and GSPC does not claim.
The 12 hazard categories (14 codes) — nearest GSPC axis, RELATED never equivalent
vcrsafetyviolent crimes
ncrsafetynon-violent crimes
srcsafetysex-related crimes
csesafetychild sexual exploitation
sshsafetysuicide & self-harm
iwpsafetyindiscriminate weapons
ipvno GSPC axisintellectual-property violations
dfmno GSPC axisdefamation
htesafetyhate
prvno GSPC axisprivacy
spc_ele / spc_fin / spc_hltno GSPC axisspecialised advice
sxc_prnsafetysexual content
8 of 14 codes sit in the harmful-request family our safety axis measures (calibrated refusal on paired requests) — a RELATED-TO join, with the caveat that refusal calibration is not a per-hazard verdict. The other 6 codes (ipv, dfm, prv, spc_ele/fin/hlt) have no GSPC axis that measures the same quantity — IP regurgitation, defamation, privacy leakage and advice quality are not refusal calibration — so no join is drawn and the gap stays visible. An earlier draft of this table (never shipped) joined all 14; review held several joins indefensible, and the deletions are the fix. A grade relative to a floating reference model is in any case not comparable across time, so grade→measurement mapping stays dishonest and undone.
The importer doctrine — bind it, do not become it
- An AILuminate result imports as an ATTACHED evidence record on the safety axis's trail — never MEASURED, never a 23rd axis, never a fused GSPC+AILuminate grade.
- Grades are recorded as the framework asserted them, with the floating-baseline caveat carried inside the record.
- Only the GHA publisher signs; the importer emits unsigned queue artefacts.