Rating the Raters — measurement, not certification
Who measures the measurers?
Benchmark operators grade the whole field and are themselves graded by nobody. This programme recomputes what a rating organisation publishes, from the rows that organisation published, using deterministic arithmetic only. No model judges anything. Where our arithmetic confirms theirs, we say so — which, on this first result, is most of it.
Read this before the number
This is result 001. One rating organisation has been measured on one axis. This is not a survey and no cross-rater comparison exists.
One rating organisation has been measured, on one axis, on one benchmark. Everything below is UNMEASURED. They are named only to state that they have NOT been measured -- no finding about any of them is expressed or implied, and their appearance here is not a ranking, a shortlist, or a queue.
What the ARC Prize project got right
ARC never claimed its 66% figure was a pass@2 score. ARC's own wording is 'average human performance', and ARC states its solvability claim -- at least 2 people in no more than 2 attempts -- separately and correctly. Nothing ARC published is false. This measurement exists because the rule-matched comparable is absent upstream, so the field routinely sets AI pass@2 scores against a human number measured under a looser rule. It is also worth saying plainly that this audit was possible ONLY because ARC publishes its participant-level rows under MIT, publishes its scoring rule verbatim, publishes gold-label corrections in a changelog rather than editing silently, and self-reports its own contamination problem. Most benchmark operators publish none of that, and cannot be audited at all. ARC's transparency is the reason it can be measured, and it should not be penalised in reputation for being the one organisation that made the check possible.
Their calibration claim: REPRODUCES
“Every ARC-AGI-2 evaluation task was solved by at least 2 people in no more than 2 attempts.” — reproduced on 161 of 161 test pairs. Reproduces on every (task, test_index) pair the published rows cover. The margin is exact rather than comfortable: at least one pair was solved by exactly 2 distinct sessions within 2 submissions, which is the claim's floor. That is a property of the claim, not a defect in it -- ARC calibrated to a threshold and the threshold is met.
Their headline figure reconciles exactly.
Exactly one of the five candidate aggregations reproduces ARC's published 66%: the macro-average over tasks under unlimited submissions, at 66.07%. ARC's figure is correct and correctly computed. No discrepancy was found in it. Their published wording: “Average human performance on these tasks in our test sample was 66%.”
The finding — a rule mismatch
The benchmark's own scoring rule, in its operator's words: “For each test input, the test-taker is allowed 2 trials. This holds for all test-takers, either humans or AI.” (github.com/arcprize/ARC-AGI-2 readme) The published human figure is computed under unlimited submissions. Under the two-trial rule the benchmark applies to machines, the comparable human number is lower:
| Scoring rule applied to humans | Human result | 95% interval |
|---|---|---|
| Unlimited submissionsthe operator's published basis | 66.07% | 61.91 – 70.03 |
| Two trialsthe benchmark's own machine rule | 55.27% | 50.97 – 59.51 |
| One trial | 40.44% | 36.12 – 44.77 |
Gap between the published figure and the rule-matched figure: 10.8 percentage points. 20.9% of human attempts used more than two submissions; 16.3% of eventually correct attempts needed more than two.
Verdict — MISMATCHED. ARC's published human figure is computed under a different attempt budget than the one ARC's scoring rule applies to AI systems, and ARC does not publish a rule-matched human figure alongside it.
The criterion, stated so it can be argued with
RTR-A1 — Human-Reference Rule Match
Is the benchmark's published human-performance figure computed under the same scoring rule the benchmark applies to machines?
Applies to: Any benchmark operator that publishes BOTH a headline human performance figure AND machine scores on the same benchmark under a stated scoring rule.
- 1. Read the operator's stated machine scoring rule. Extract the attempt budget k (trials permitted per test item) and the grading function.
- 2. Read the operator's published human figure and the computation the operator states for it.
- 3. If the operator publishes participant-level rows, recompute the human figure under the machine rule -- same k, same grading, same aggregation -- to obtain H_matched.
- 4. Report H_published, H_matched, and gap = H_published - H_matched in percentage points.
MATCHED
H_published is computed at the same k and grading as machine scores. Gap is 0 by construction.
MISMATCHED
H_published is computed at a different k or grading, and the operator publishes no rule-matched human figure. Gap reported in pp.
NOT_APPLICABLE
The operator publishes no headline human figure.
UNMEASURABLE_FROM_PUBLIC_ROWS
The operator publishes a human figure but not the rows needed to recompute it. This is a limit on what CSOAI can check, not a finding against the operator.
- · Deterministic arithmetic only. No model judges any output.
- · Grading is the operator's own grading function, unmodified.
- · n >= 30 on the aggregation unit or the result is not published.
- · Every published number recomputable from published rows.
Disagree with this? The criterion is stated so it can be argued with. An operator disputing this result should name which step it rejects: the extracted budget k, the aggregation, the stop-at-first-correct assumption, or the claim that no rule-matched figure is published. Send to [email protected]; disputes are published verbatim alongside the result.
The assumption this rests on
A participant stopped submitting once correct, so a solved attempt's correct submission is its last submission.
ARC publishes submission counts, not the index of the correct submission. Without this assumption an attempt recorded as '5 submissions, 1 correct' cannot be placed inside or outside a 2-submission budget.
Evidence for it, computed from the same rows: 772 of 773 solved attempts have exactly one correct submission, and the submission-count distribution over solved attempts is monotonically decreasing (472, 175, 70, 30, 16, 6, 3, 1) — the shape produced by stopping at first success. Forcing the single anomalous row to count the other way moves the headline to 55.34%.
It remains an assumption, not a fact. Publish the 1-based index of the correct submission as a column. One column removes the assumption entirely.
Why the interval is a bootstrap and not Wilson
Our house standard is the Wilson score interval, and it does not apply to this headline. The headline is a mean of per-task proportions, not a single binomial proportion. We publish all three so the difference is visible rather than asserted.
| Bootstrap over tasks — published | 50.97 – 59.51 |
| Wilson misapplied to the per-task mean — not used | 46.54 – 64.4 |
| Attempt-level rate 51.35% (647/1260) — Wilson, where it is exact | 48.59 – 54.1 |
percentile bootstrap over the 115 tasks, 10000 resamples, seed 20260826. Wilson score interval; applies exactly, the attempt-level rate is a single binomial proportion. Attempts are clustered within sessions and tasks, so nominal coverage is optimistic.
What this does not measure
- · Scoped to the 115 of 120 public-eval tasks the published rows cover. Not a statement about the other 5.
- · Human rows are public-eval; the leaderboard's AI scores are semi-private. The two populations are different task sets, so the human and AI figures are not directly comparable even after the rule is matched. This result narrows one of the two mismatches, not both.
- · The pass@k derivation rests on the stop-at-first-correct assumption documented in assumption_check.
- · Human participants were not incentivised, timed, or selected the way an evaluated model is prompted. This measures the published figures' rule-consistency, not human ability.
Also unmeasured, about this benchmark specifically:
- · ARC-AGI-2 semi-private and private evaluation results. These are not third-party recomputable by design, for a defensible anti-contamination reason. CSOAI cites them, never restates them as CSOAI-measured.
- · ARC-AGI-1 and ARC-AGI-3 on this axis.
- · The 5 public-eval tasks the published rows do not cover.
- · Whether ARC's grading of model outputs is correct. That is a separate measurement and has not been done.
Coverage. A limit on what public data can settle, not a defect in ARC. ARC's own dataset card states the released rows are not comprehensive. Every figure here is scoped to the 115 covered tasks and none of them should be read as covering the other 5. Covered: 115 of 120 public-eval tasks (95.8%). The remaining 5 are UNMEASURED.
The rows — recheck us
The 115 rows every headline number is computed from. The macro figure for a rule is the unweighted mean of that rule's rate column. Measured from arcprize/arc_agi_2_human_testing (test_pair_attempts.csv, MIT), SHA-256 d13f1a05580285a04f6bba4dafefdc4fd81a45014c134f6625d1ab625d7deea5. Benchmark pinned to commit f3283f727488ad98fe575ea6a5ac981e4a188e49. ARC corrected public-eval gold labels 20+ times between 2025-03-24 and 2025-04-17, published in a changelog rather than applied silently. An unpinned ARC number is not meaningful. Recompute with scripts/rating_the_raters_001_arc.py.
n = 115 tasks · 161 test pairs · 1260 attempts · 392 sessions. Our floor for publication is 30.
Result RTR-001 · subject: ARC Prize -- ARC-AGI-2 public evaluation human baseline · measured 2026-08-26. Nothing on this page is signed, and it confers no status on the organisation measured. Council of AI recomputes published claims and reports what reproduced; it is not a certification body and issues no approval. Their semi-private results are cited where relevant and are never restated as measured by us.
A note on where this came from. An internal strategy draft once proposed publishing a figure for how many rating organisations keep a corrections record. That figure had never been measured and it is not published here. What is published is one recomputation, of one claim, by one organisation, with its rows attached — which is the only kind of thing this programme will ever publish.