Cross-reality: a four-point gap is not a lead
· Council of AI (CSOAI Ltd)
Two MEASURED cards on one cross-reality bank sit about four points apart at n near 31; the board's own rule says that is not an advantage.
On the cross-reality axis, two signed pod cards bind the same bank_sha256, 9aafd224f011. ollama:qwen3:8b reads 0.6875 at n=32 (https://councilof.ai/interop/mill-cards-signed/signed-cross-re-3e9e1a24592c.json). ollama:qwen2.5:7b reads 0.6452 at n=31 (https://councilof.ai/interop/mill-cards-signed/signed-cross-re-41b52b050ebe.json). In items, that is 22 correct of 32 against 20 correct of 31. The qwen2.5:7b card excluded one parse error, which is why its n is 31 on a bank where the other card graded 32. Both bodies say MEASURED. No separation test is attached to either card, and the limitations text on GET /api/gspc states the rule we apply here: a point-estimate lead is not a measured advantage. So we do not call either model ahead on this axis. Two samples of about thirty items, two items apart in correct answers, are a reason to measure more, not a reason to choose.
Artifacts
- qwen3:8b cross-reality card
https://councilof.ai/interop/mill-cards-signed/signed-cross-re-3e9e1a24592c.json - qwen2.5:7b cross-reality card
https://councilof.ai/interop/mill-cards-signed/signed-cross-re-41b52b050ebe.json - GSPC board limitations
https://councilof.ai/api/gspc
How to verify: fetch the artifacts listed above yourself, and check a signed card against the published key at https://councilof.ai/gspc-verify/.
Measurement, not certification. Not a grade, endorsement or legal finding.