Every model below was asked the same 918 verified factual questions cold: no tools, no web access, no retrieval, instructed to say "not sure" rather than guess. Answers were graded against ground truth that we independently re-derived from live primary sources (openFDA, SEC EDGAR, USPTO, ClinicalTrials.gov, Companies House, Wikidata, NIST, real code execution) and adversarially audited three times.
| Model | Correct | Hallucinated | Hedged honestly |
|---|---|---|---|
| Claude Opus tier | 81.1% | 4.8% | 14.2% |
| Claude Sonnet 5 | 70.8% | 10.1% | 18.6% |
| Claude Haiku 4.5 | 65.7% | 11.5% | 22.8% |
Read it plainly: even the most capable tier confidently invents roughly one answer in twenty when it can't look things up, and the small fast tier one in nine. Capability helps, but it doesn't make the problem disappear, and on live-data domains (recalls, filings, registries) every tier gets materially worse. That gap is exactly what these datasets measure.
| Dataset | Haiku 4.5 | Sonnet 5 | Opus tier |
|---|---|---|---|
| FDA Drug/Device Safety | 27.5% | 14.3% | 10.0% |
| Citation-Graph Verification | 27.9% | 32.6% | 4.7% |
| Clinical Trial Outcomes | 23.4% | 19.1% | 10.6% |
| SEC EDGAR Financials | 17.1% | 7.3% | 4.9% |
| Language & Runtime Semantics | 15.4% | 5.1% | 5.1% |
| UK Companies House | 9.1% | 1.8% | 12.7% |
| Cross-Source Verification | 9.5% | 14.3% | 11.9% |
| Scientific Claims | 11.6% | 11.6% | 7.0% |
| Fact-Verification (general) | 9.7% | 6.8% | 0.8% |
| English Language & Attribution | 9.5% | 14.3% | n/a* |
| Geographic Facts | 7.5% | 6.7% | 2.5% |
| Government Contracts | 7.5% | 7.5% | 5.0% |
| Legal Citation & Case Law | 7.1% | 5.7% | 7.1% |
| Code-Reality Verification | 7.1% | 11.9% | 4.8% |
| Mathematical Claims | 7.0% | 7.0% | 0.0% |
| Patent & IP Claims | 2.5% | 20.0% | 5.0% |
| Historical Facts | 4.8% | 0.0% | 0.0% |
*Opus run on the English-literature set was blocked by a content filter while reproducing famous quotations; it will be re-run. Methodology: each model answered from its own knowledge only (agents instructed to use no tools, given a questions-only file that contains no answers); a separate grader model compared every answer against the verified ground truth as correct / hallucinated / hedged. Sonnet 5 figures are from the per-row model-graded layer that ships inside every dataset, so you can audit its grading row by row. Run dates: 2026-07-11 (Sonnet), 2026-07-17 (Haiku, Opus).
Want your model on this table, or your own domain measured? Run the bench yourself with any model or stack, or ask about a Hallucination Report for documented, hand-off-able accuracy evidence.