Governance: two models, one frozen bank, a gap in point estimates

· Council of AI (CSOAI Ltd)

Two MEASURED pod cards bind the same governance bank hash; their accuracies differ and no separation test is attached.

Two signed pod cards on the governance axis bind the same bank_sha256, b93f9808f014. ollama:qwen3:8b reads accuracy 0.5865 at n=237 (https://councilof.ai/interop/mill-cards-signed/signed-governan-09759e29f44b.json). ollama:gemma3:4b reads accuracy 0.2785 at n=237 (https://councilof.ai/interop/mill-cards-signed/signed-governan-6182304443ff.json). Both card bodies say MEASURED, and both record zero transport errors and zero parse errors excluded. The two cards carry different instrument_sha256 values, so the shared thing a reader can confirm is the bank, not an identical instrument. Neither card carries a separation test. Read this as a gap between two point estimates on one frozen bank, not as a ranking of the two models. The governance axis entry on GET /api/gspc describes the task as EU AI Act risk-tier classification. These pod cards commit to an items_sha256 but do not link the item rows, so a stranger can check who signed which numbers, but cannot re-grade the items from these two cards alone.

Artifacts

How to verify: fetch the artifacts listed above yourself, and check a signed card against the published key at https://councilof.ai/gspc-verify/.

Measurement, not certification. Not a grade, endorsement or legal finding.

All evidence notes · /feeds/notes.xml