Art5-safeguard: a wide gap, stated as two point estimates
· Council of AI (CSOAI Ltd)
Two MEASURED cards on the same art5-safeguard bank differ widely; without a separation test it is still not a ranking.
On the art5-safeguard axis, two signed pod cards bind the same bank_sha256, bdb7c0e8b7ad. ollama:qwen3:8b reads 0.9722 at n=36 (https://councilof.ai/interop/mill-cards-signed/signed-art5-saf-b15988212458.json). ollama:qwen2.5:0.5b-instruct reads 0.5278 at n=36 (https://councilof.ai/interop/mill-cards-signed/signed-art5-saf-6bd698f292e8.json). In items, that is 35 correct of 36 against 19 correct of 36. Both bodies say MEASURED with zero parse errors excluded. The gap is wide for a 36-item bank, but neither card carries a separation test, so we state it as two point estimates and not as a verdict on either model. A card on this axis is a measurement against one frozen bank; it is not a legal finding about any deployment of either model. To see where other models in the pod population landed on the same bank, filter GET /api/hub-cards by axis and follow each card URL listed there, checking that every card you compare names the same bank_sha256.
Artifacts
- qwen3:8b art5-safeguard card
https://councilof.ai/interop/mill-cards-signed/signed-art5-saf-b15988212458.json - qwen2.5:0.5b-instruct art5-safeguard card
https://councilof.ai/interop/mill-cards-signed/signed-art5-saf-6bd698f292e8.json - Hub cards endpoint
https://councilof.ai/api/hub-cards
How to verify: fetch the artifacts listed above yourself, and check a signed card against the published key at https://councilof.ai/gspc-verify/.
Measurement, not certification. Not a grade, endorsement or legal finding.