Council of AI · evidence surface · /genai-mil
Three vendors now answer to millions of federal seats. Who measured their behaviour?
On 31 August 2026 the U.S. War Department put three frontier AI vendors on one platform, GenAI.mil, at Impact Level 5, for a user base designed to reach more than three million people. The security wrapper is authorised. The deployments are sealed. The models' behaviour on the governance-and-safety axes is unsigned. That last gap is the thing this page is about.
The rule this whole page keeps: we measure; we do not certify. Nothing here is a conformity mark, an endorsement, or a legal determination. A signed card records what a reproducible test found. Where we did not test, we say UNMEASURED. Where a system is sealed from the public and cannot be tested, we say UNCHECKABLE — and that is an honest, first-class answer, not a failure.
What was announced (the public facts, signed)
These are recorded as public-notice cards: we sign the press, not the models. Each hashes an official or reported source and states plainly what it does not measure.
| Fact | Value | Source |
|---|---|---|
| Vendors on GenAI.mil | OpenAI ChatGPT (ChatGPT Mil), xAI Grok (Grok for Government / Starshield AI), Google Gemini | War Dept release; DefenseScoop; Military Times; TechCrunch |
| Accreditation | Impact Level 5 (IL5) — highest for non-public, sensitive unclassified data | War Dept release |
| Design capacity | More than 3 million military members, civilian employees and contractors | War Dept release |
| Users on the platform | ~1.5 million (as of 12 Jun 2026, Pentagon CTO) | Military Times / DefenseScoop |
Note on counts: the sourced current-user figure is ~1.5M (June 2026). A ~1.7M figure circulated in operator notes is unverified; we record only the sourced number.
The three states, honestly
UNCHECKABLE The military deployments
ChatGPT Mil and Grok for Government run inside an IL5 boundary that is not public. We did not probe them, and we will not: a classified or government instance is UNCHECKABLE by definition. Any claim that we measured "the model the military uses" would be false. We measured the public model, never the sealed instance.
AUTHORISED The FedRAMP wrapper — and what it does not cover
The security wrapper around these products is authorised on the FedRAMP marketplace: OpenAI ChatGPT Enterprise and API (marketplace listing FR2533155773, ongoing as of 9 Jan 2026); Google Gemini at FedRAMP High (Mar 2025); Anthropic Claude at FedRAMP High via AWS and Google Cloud (Apr–Jun 2025). FedRAMP authorises cloud security controls. It does not sign the model's behaviour on the behavioural axes. That gap is us.
UNMEASURED The public models' behaviour — queued, not done
The honest state today: Council of AI has no signed behavioural card yet for the public GPT, Grok, Gemini, or Claude models. Grading the public models on the 14 behavioural GSPC axes with the deterministic grader is queued as an owner step (it needs the frontier-inference key). Until that runs, this row stays UNMEASURED. We will not print a score we did not measure.
Vendor claims vs our card
Each vendor published a system or model card documenting its own safety testing and known limitations — public disclosures, on the record. We record the disclosure and, beside it, our own state. Public disclosures compared, not certified; factual, not defamatory.
| Vendor | Public disclosure | Our card |
|---|---|---|
| OpenAI (GPT / ChatGPT) | GPT-5 System Card | UNMEASURED |
| Anthropic (Claude) | Claude Opus system card | UNMEASURED |
| Google (Gemini) | Gemini model cards | UNMEASURED |
| xAI (Grok) | Grok model card | UNMEASURED |
A vendor's self-reported testing is a claim. An independent, reproducible, signed behavioural card is a measurement. The point of the comparison is the shape of the gap between the two — not to dispute what any vendor published.
Where the frameworks meet the axes
On 18 March 2026 CAISI (the NIST Center for AI Standards and Innovation) signed an MOU with GSA to bring AI-evaluation science into federal procurement through the USAi platform. NIST's AI Risk Management Framework is the federal risk spine. Our crosswalk points each RMF function at the GSPC behavioural axis that produces signed, reproducible evidence for it — NIST AI RMF → GSPC axes crosswalk. It is a mapping, not a legal determination.
Open door: the GSA contract-clause comment window (GSAR 552.239-7001) closed on 3 August 2026 — that door is shut. But NIST's Trustworthy AI in Critical Infrastructure Profile Community of Interest is open now, by mailing list and community channel. That is where an independent measurement layer belongs, and where we are preparing a contribution.
The compute thesis
A cluster does not produce a signed card.
The same week these deployments went live, Anthropic was reported to have signed a roughly 35-billion-dollar cloud-compute deal with Nvidia-backed Lambda — one of several very large capacity commitments across the frontier labs. Compute buys capability. It does not, on its own, produce a reproducible, independently signed record of how a model behaves on the axes that a government cares about. That record has to be measured. Measuring it — for everyone, in the open — is the whole job.
Verify it yourself
- Live board and axis totals: /api/gspc (totals.public_count is the sentence to quote).
- Signed public root (Merkle over every published card): /root.json.
- Signed card index and per-card verifier: /signed/.
- The GenAI.mil facts above are published as public-notice cards folded into the signed root; the frontier-model behavioural cards will appear on the leaderboard once the owner-step grading runs.
Council of AI measures AI behaviour and publishes signed, reproducible records. We do not certify, accredit, or endorse, and verification is free. The GenAI.mil deployments are non-public and were not probed; the public models named here are not yet measured. This page states facts and states plainly what remains UNCHECKABLE or UNMEASURED. Last updated 1 September 2026.