Care: read n and the exclusion count before the accuracy

· Council of AI (CSOAI Ltd)

On one care bank, the higher accuracy sits on fewer graded items, because most of that model's replies could not be parsed.

Two signed pod cards on the care axis bind the same bank_sha256, cee5e47a31f1, but they did not grade the same number of items. ollama:llama3.2:3b reads 0.7143 at n=77 (https://councilof.ai/interop/mill-cards-signed/signed-care-83d57579098c.json), and its compute_evidence records parse_errors_excluded: 122. ollama:qwen3:8b reads 0.3434 at n=198 (https://councilof.ai/interop/mill-cards-signed/signed-care-4a4ede4e8c70.json), with parse_errors_excluded: 1. The higher accuracy therefore sits on the replies that could be parsed, and on the llama3.2:3b card those were fewer than the replies that could not. Both bodies say MEASURED, because both n values clear 30. A reader comparing the two should read n and the exclusion count before the accuracy: a model that replies out of format on most items and correctly on the rest is a different finding from a model that replies in format on nearly every item. Neither card carries a separation test.

Artifacts

How to verify: fetch the artifacts listed above yourself, and check a signed card against the published key at https://councilof.ai/gspc-verify/.

Measurement, not certification. Not a grade, endorsement or legal finding.

All evidence notes · /feeds/notes.xml