CSOAI - for healthcare + life sciences
Clinical AI, measured - and the measurement signed
Clinical AI is squarely high-risk under the EU AI Act, on top of MDR/IVDR and, in the US, HIPAA. CSOAI measures how models behave on frozen, published banks and signs each run. We do not assess your device, and we issue nothing.
Your path with CSOAI
What you can check, right now
Clinical AI fails in ways a generic benchmark does not look for: whether the model gives calibrated care rather than confident care, whether it refuses a harmful instruction, whether it stays inside the Article 5 prohibitions, and how it handles a distressed user. Every row carries its own n — CareBench is the largest bank on the board and Art5Bench is one of the smallest.
| Axis | Bench | n | Leader accuracy | 95% CI | Separation |
|---|---|---|---|---|---|
| care | CareBench | 199 | 53.5% | 46.6–60.3% | SEPARATED p=0.0356 |
| safety | DefBench | 36 | 94.4% | 81.9–98.5% | TIE — indistinguishable p=0.6875 |
| art5-safeguard | Art5Bench | 36 | 97.2% | 85.8–99.5% | TIE — indistinguishable p=1 |
| affect | AffectBench | 41 | 87.8% | 74.5–94.7% | SEPARATED p=0.0078 |
high-risk
Diagnosis and triage qualify - by two different routes
AI that is a safety component of a device regulated under MDR or IVDR is high-risk through Article 6(1) and Annex I. Triage of patients in emergency healthcare is listed separately in Annex III. Same tier, different conformity paths - and the route decides who assesses you.
What you can check on the board
Two behavioural axes speak to clinical conduct — whether a model refuses in a calibrated way, and what it costs in care terms when protecting and helping pull against each other.
Every figure below is read live from GET /api/gspc when this page loads — no count is typed into it. These axes measure how a fleet of models behaves on a frozen, published bank on a date. None of them is an assessment of your organisation, and none is a conformity opinion: determination stays with your regulator.
Board right now: 22 axes · 15 measured — a published slot is not a measurement, which is why both numbers travel together.
Care-cost under paired conduct scenarios: protecting and helping traded off against each other, which is the shape of most clinical edge cases.
- What is graded:
- care-cost (protect × help) under paired conduct scenarios
- n:
- 199 bank items
- Leader:
- 53.5% (council-ethics-v3-light (council specialist))
- Separation:
- SEPARATED
- Frozen bank:
- csoai/gspc-care
Calibrated refusal on paired requests — refusing the harmful sibling of a request without refusing the benign one. Over-refusal is a clinical failure too.
- What is graded:
- calibrated refusal on paired requests
- n:
- 36 bank items
- Leader:
- 94.4% (gemma3:12b (base model))
- Separation:
- TIE — the leader's lead is not statistically separated, and a tie is never counted as a win.
- Frozen bank:
- csoai/gspc-agi
Emotional and embodied safety: manipulation, disclosure and handling of vulnerability. Its legal gold labels are counsel-pending, which the board states on the row.
- What is graded:
- emotional & embodied safety (manipulation / disclosure / vulnerability)
- n:
- 41 bank items
- Leader:
- 87.8% (council-preservation-v3-light (council specialist))
- Separation:
- SEPARATED
- Frozen bank:
- csoai/gspc-affect
Re-check it without us
Each published measurement card is Ed25519-signed over its exact bytes, and its id is the sha256 of those bytes. Pin the key from our DID document first — a card verified against the key it ships with proves only that the file is self-consistent.
- /signed/HOW-TO-VERIFY.md — the four commands, start here
- /signed/card_index.json — the signed index of published cards
- /.well-known/did.json — the key to pin against
- /gspc-verify — recompute the replay chain in your browser, no account
Questions, answered
Yes. AI that is a safety component of a product covered by the device legislation in Annex I - MDR and IVDR - is high-risk under Article 6(1), and AI used for triage of patients in emergency healthcare is separately listed in Annex III. Both trigger conformity, oversight and documentation duties.
No - it applies alongside them. Documentation can be aligned into a single technical file, but the AI-specific obligations are additional.
Answers from published measurement, or it refuses. Your question is typed into the lobby — nothing sends until you press Ask.
Deterministic pane commands · grounded /api/chat lane · consent checkpoint on consequential steps