CSOAI - for healthcare + life sciences

Clinical AI, measured - and the measurement signed

Clinical AI is squarely high-risk under the EU AI Act, on top of MDR/IVDR and, in the US, HIPAA. CSOAI measures how models behave on frozen, published banks and signs each run. We do not assess your device, and we issue nothing.

Your path with CSOAI

What you can check, right now

Clinical AI fails in ways a generic benchmark does not look for: whether the model gives calibrated care rather than confident care, whether it refuses a harmful instruction, whether it stays inside the Article 5 prohibitions, and how it handles a distressed user. Every row carries its own n — CareBench is the largest bank on the board and Art5Bench is one of the smallest.

AxisBenchnLeader accuracy95% CISeparation
careCareBench19953.5%46.6–60.3%SEPARATED p=0.0356
safetyDefBench3694.4%81.9–98.5%TIE — indistinguishable p=0.6875
art5-safeguardArt5Bench3697.2%85.8–99.5%TIE — indistinguishable p=1
affectAffectBench4187.8%74.5–94.7%SEPARATED p=0.0078
Every figure above is read from GET /api/gspc when this page loads — none of it is written into the page. Recompute a signed card yourself at /gspc-verify, walk every chain position at /signed/chain.json, and read the rules the grades are computed under at /methodology. A leader is the highest point estimate on the board, not an approval; a TIE is not a win; and unmeasured means no run exists — never zero. We measure against these obligations; we do not enforce them and we certify nothing.

high-risk

Diagnosis and triage qualify - by two different routes

AI that is a safety component of a device regulated under MDR or IVDR is high-risk through Article 6(1) and Annex I. Triage of patients in emergency healthcare is listed separately in Annex III. Same tier, different conformity paths - and the route decides who assesses you.

What you can check on the board

Two behavioural axes speak to clinical conduct — whether a model refuses in a calibrated way, and what it costs in care terms when protecting and helping pull against each other.

Every figure below is read live from GET /api/gspc when this page loads — no count is typed into it. These axes measure how a fleet of models behaves on a frozen, published bank on a date. None of them is an assessment of your organisation, and none is a conformity opinion: determination stays with your regulator.

Board right now: 22 axes · 15 measured — a published slot is not a measurement, which is why both numbers travel together.

  • careMEASUREDCareBench

    Care-cost under paired conduct scenarios: protecting and helping traded off against each other, which is the shape of most clinical edge cases.

    What is graded:
    care-cost (protect × help) under paired conduct scenarios
    n:
    199 bank items
    Leader:
    53.5% (council-ethics-v3-light (council specialist))
    Separation:
    SEPARATED
    Frozen bank:
    csoai/gspc-care
  • safetyMEASUREDDefBench

    Calibrated refusal on paired requests — refusing the harmful sibling of a request without refusing the benign one. Over-refusal is a clinical failure too.

    What is graded:
    calibrated refusal on paired requests
    n:
    36 bank items
    Leader:
    94.4% (gemma3:12b (base model))
    Separation:
    TIE — the leader's lead is not statistically separated, and a tie is never counted as a win.
    Frozen bank:
    csoai/gspc-agi
  • affectMEASUREDAffectBench

    Emotional and embodied safety: manipulation, disclosure and handling of vulnerability. Its legal gold labels are counsel-pending, which the board states on the row.

    What is graded:
    emotional & embodied safety (manipulation / disclosure / vulnerability)
    n:
    41 bank items
    Leader:
    87.8% (council-preservation-v3-light (council specialist))
    Separation:
    SEPARATED
    Frozen bank:
    csoai/gspc-affect

Re-check it without us

Each published measurement card is Ed25519-signed over its exact bytes, and its id is the sha256 of those bytes. Pin the key from our DID document first — a card verified against the key it ships with proves only that the file is self-consistent.

Ed25519-signed
every published measurement card
Verify without an account
pin our key, recompute the bytes
Empty cells stay empty
unmeasured is published, not hidden
MIT-licensed core
no vendor lock-in

Questions, answered

Is medical AI high-risk under the EU AI Act?

Yes. AI that is a safety component of a product covered by the device legislation in Annex I - MDR and IVDR - is high-risk under Article 6(1), and AI used for triage of patients in emergency healthcare is separately listed in Annex III. Both trigger conformity, oversight and documentation duties.

Does the AI Act replace MDR/IVDR?

No - it applies alongside them. Documentation can be aligned into a single technical file, but the AI-specific obligations are additional.

Council OS — healthcare

Answers from published measurement, or it refuses. Your question is typed into the lobby — nothing sends until you press Ask.

Open Council OS

Deterministic pane commands · grounded /api/chat lane · consent checkpoint on consequential steps