Leaderboard

How much do frontier models hallucinate? We measured it.

Every model below was asked the same 918 verified factual questions cold: no tools, no web access, no retrieval, instructed to say "not sure" rather than guess. Answers were graded against ground truth that we independently re-derived from live primary sources (openFDA, SEC EDGAR, USPTO, ClinicalTrials.gov, Companies House, Wikidata, NIST, real code execution) and adversarially audited three times.

Overall, cold (no tools)

ModelCorrectHallucinatedHedged honestly
Claude Opus tier81.1%4.8%14.2%
Claude Sonnet 570.8%10.1%18.6%
Claude Haiku 4.565.7%11.5%22.8%

Read it plainly: even the most capable tier confidently invents roughly one answer in twenty when it can't look things up, and the small fast tier one in nine. Capability helps, but it doesn't make the problem disappear, and on live-data domains (recalls, filings, registries) every tier gets materially worse. That gap is exactly what these datasets measure.

Cold hallucination rate by domain

DatasetHaiku 4.5Sonnet 5Opus tier
FDA Drug/Device Safety27.5%14.3%10.0%
Citation-Graph Verification27.9%32.6%4.7%
Clinical Trial Outcomes23.4%19.1%10.6%
SEC EDGAR Financials17.1%7.3%4.9%
Language & Runtime Semantics15.4%5.1%5.1%
UK Companies House9.1%1.8%12.7%
Cross-Source Verification9.5%14.3%11.9%
Scientific Claims11.6%11.6%7.0%
Fact-Verification (general)9.7%6.8%0.8%
English Language & Attribution9.5%14.3%n/a*
Geographic Facts7.5%6.7%2.5%
Government Contracts7.5%7.5%5.0%
Legal Citation & Case Law7.1%5.7%7.1%
Code-Reality Verification7.1%11.9%4.8%
Mathematical Claims7.0%7.0%0.0%
Patent & IP Claims2.5%20.0%5.0%
Historical Facts4.8%0.0%0.0%

*Opus run on the English-literature set was blocked by a content filter while reproducing famous quotations; it will be re-run. Methodology: each model answered from its own knowledge only (agents instructed to use no tools, given a questions-only file that contains no answers); a separate grader model compared every answer against the verified ground truth as correct / hallucinated / hedged. Sonnet 5 figures are from the per-row model-graded layer that ships inside every dataset, so you can audit its grading row by row. Run dates: 2026-07-11 (Sonnet), 2026-07-17 (Haiku, Opus).

Want your model on this table, or your own domain measured? Run the bench yourself with any model or stack, or ask about a Hallucination Report for documented, hand-off-able accuracy evidence.