Datasets

Fixed-scope datasets for verification, eval, and agentic training

Smaller than the full landscape dataset, and priced for it — each one is a real, verified, one-time snapshot. No estimated fields, no synthetic trajectories, no forced true/false calls where the honest answer is "unverifiable."

Mathematical Claims Verification

Theorem, constant, and attribution claims checked by real computation, not recall

$349 once

Forty-three real mathematical claims (famous theorem statements, numeric sequence/constant properties, computational results, open-problem status checks, and "who first proved X" attributions), each checked against code actually written and run (Python sympy/mpmath, a from-scratch BigInt library, and an exhaustive brute-force combinatorial search), the real Wikidata knowledge graph, real published papers via OpenAlex, and, for one live current-events check, primary web sources fetched live. An AI model asked to check its own math claim just restates it more confidently; this dataset actually computes, queries, and looks the answer up, and says so honestly on the rare row where a second independent check genuinely doesn't exist. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Every claim also tested live against Claude Sonnet 5 asked cold, with no source access: 3 of 43 (7%) answers were confidently wrong, including misattributing 2^31− 1's primality proof to Édouard Lucas instead of the verified Euler
  • Re-audited every original row live and found one real self-correction: Wikidata does carry a discoverer statement for 2^31−1 being prime (Euler, 1772), on the specific-number entity, not the generic Mersenne-prime class entity checked the first time
  • Added an exhaustive brute-force proof that the Ramsey number R(3,3)=6 (all 32,768 edge-colorings of K6 checked by real code, every one contains a monochromatic triangle), and confirmed Wikidata's "credits the poser, not the solver" pattern on 3 more rows
  • Live-fetched today's actual largest known prime (2^136,279,841−1, GIMPS, Oct 2024) from two independent primary sources and found Wikidata's own record for the same fact frozen at a 1951–52 value — a real, current structured-data staleness gap

Scientific Claims Verification

Physical constants and discovery claims checked against NIST, OpenAlex, Wikidata, and real independent computation

$349 once

Forty-three real claims about physical constants, first-discovery/first-measurement attributions, and specific published experimental results across physics, chemistry, astronomy, biology, and earth science, each checked against NIST's live CODATA fundamental-constants database, the actual paper reporting a result via OpenAlex, Wikidata's structured discoverer/date/value statements, or, new this pass, real independently executed computation deriving one fundamental constant from others. Fills the gap between Groundtruth's clinical-trial and citation-graph products, which don't touch general physical-science claims, and is now the catalog's most methodologically diverse verification product. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Every claim also tested live against Claude Sonnet 5 asked cold: 5 of 43 (12%) answers were confidently wrong, including fabricating a different 2015 Avogadro-constant figure than the real IAC-published value
  • Ran two from-scratch computational cross-checks this session, independently re-deriving the Rydberg constant and the c=1/√(μ₀ε₀) identity from five raw CODATA inputs each via real executed Python, agreeing with NIST's own published values to 12-13 significant figures
  • Found Wikidata's own unreferenced discovery-year statement for buckminsterfullerene (1984) contradicts the actual discovery paper's real publication date (November 1985), and found four disagreeing referenced river-length sources on Wikidata's own Amazon entity spanning nearly 600 km
  • Re-verified all 24 original rows live with zero drift, then added a genuinely independent second method to 7 rows, while explicitly declining to force a weak second-method addition where the only matching source didn't actually say what it needed to say

Patent & IP Claims Verification

Patent and IP claims checked against the live USPTO record, not what a company says about them

$309 once

Forty real claims companies and named inventors have made about patents ("we hold a patent on X," "patent number Y covers Z," "first to patent W"), each checked against the actual live record at USPTO's Patent Public Search system: does the patent exist, is it assigned to the claimed entity, do its real granted claims actually cover what's being claimed, and is it still valid today. Every row is now also cross-checked against a second, independent USPTO data feed — the live Patent Assignment and PTAB docket history tracking what actually happened to a patent after it was granted, not just its bibliographic snapshot — which caught patents that quietly lapsed for non-payment years after being cited in litigation, ownership chains that don't match a company's own marketing page, and three real errors in this dataset's own earlier release, corrected in place and disclosed rather than silently fixed. The original 23 rows also carry a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Expanded from 23 to 40 rows with a second independent verification method added to every row — USPTO's own live Patent Assignment and PTAB docket feed, not just a bibliographic lookup — which caught three real errors in the original release's own explanations and disclosed all three in place rather than silently rewriting them
  • Found a company's live product page still claiming present-tense patent protection under a specific patent number that USPTO's own record shows actually lapsed for non-payment more than five years ago
  • Traced Theranos's real patent portfolio surviving Elizabeth Holmes's fraud collapse: live USPTO assignment records independently confirm it was sold to a Fortress Investment Group shell company, which later used the same patents to sue an unrelated COVID-19 test maker

Clinical Trial Outcomes Verification

Does the trial data behind a drug claim actually back it up — checked against the real registry

$319 once

Forty-seven real claims about clinical trial outcomes and FDA drug approvals — from company press releases, SEC filings, FDA announcements, and news coverage — each checked against the actual ClinicalTrials.gov registration and results record, openFDA's real approval data, FDA's Drugs@FDA and Purple Book databases, and, where a specific paper is claimed, the real linked publication via PubMed/OpenAlex. Surfaces genuine, checkable gaps and traps: several FDA-approved, NEJM-published drugs — including the first-ever CRISPR therapy — have zero structured results posted on their own government registry entry; two different drugs (bardoxolone methyl and sotorasib) met their trial's own statistically significant primary endpoint and were still rejected or held back by FDA anyway; and two real 2023 drug-label expansions (Trikafta, Vyvgart) turned out to require an entirely new FDA application number rather than a supplement — a trap the model falls into when asked cold. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Every claim also tested live against Claude Sonnet 5 asked cold: 9 of 47 (19%) answers were confidently wrong, including two cases where the model correctly recalled a real drug approval but wrongly guessed it was filed as a supplement rather than an entirely separate FDA application
  • Two different drugs — bardoxolone methyl and sotorasib (Lumakras) — met their confirmatory trial's own statistically significant primary endpoint and were still rejected or held back by FDA, confirmed live against both the trial registry's structured statistics and openFDA's actual submission history
  • 43 of 47 rows now carry two or more independent live-checked sources per claim (trial registry, FDA approval record, a second FDA database, and/or the actual publication) — including a real transcription error this pass found in a prior release's Aduhelm/aducanumab row and fixed with the correction disclosed in place

Government Contract Verification

Federal contracting claims checked against the actual USASpending.gov award record

$299 once

Forty real claims companies make about their US federal and defense contracting relationships — "prime DoD contractor," specific dollar figures, named agency relationships — each checked against the real award data on USASpending.gov, the federal government's own system of record for who it has actually paid, with a second independent method (real SEC EDGAR filings) added wherever one genuinely exists. Surfaces real, checkable failure modes: contract-ceiling figures presented as if already earned, a real NYSE-listed company with billions in current federal work that's entirely invisible under its own name because every dollar is still filed under its pre-merger legal entity, and — new this pass — a household-name defense prime whose federal business is invisible to a simple name search because the identical legal entity is registered under more than 100 separate, un-consolidated USASpending IDs. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Added a second, independent verification method (real SEC EDGAR filings, not just USASpending) to 8 rows — e.g. quantifying that CrowdStrike's audited $4.8B SEC-reported revenue is over 24,000x its entire confirmed direct-prime federal award history, and independently confirming via a wholly different accounting concept that 83.6% of Lockheed Martin's audited net sales trace to federal contracts
  • A live re-audit of the original 25 rows found zero errors but surfaced a brand-new failure mode: Northrop Grumman's own legal name is registered as 158 separate, un-consolidated USASpending recipient IDs, undercounting its real $42B-revenue federal business by nearly 70% under a naive name search
  • Every claim also tested live against Claude Sonnet 5 asked cold: only 5 of 40 answered correctly with real confidence, 3 were confidently wrong (hallucinated), and 31 honestly hedged rather than guessing at specific post-cutoff contract figures

Web-Navigation Trajectories

Real browser sessions with captured structural data, including real failures

$249 once

Forty-two real browser-use sessions — an AI system actually navigating, clicking, typing, hovering, and scrolling a live browser against real public websites, with the exact action sequence, the real accessibility-tree/selector data captured at each step, and the real fact or outcome found at the end — plus, on a dozen rows, a second genuinely independent verification method (Wikidata's structured API, live REST API calls, raw GitHub source, NIH PubChem, caniuse.com) layered on top of the original browser session. Includes multi-hop tasks with genuine runtime branching, multi-tab flows, hover-triggered menus, infinite scroll, a fully unplanned Special:Random navigation whose destination was unknown before the session ran, and five honestly-recorded failures and partial-failures (CAPTCHA walls, login walls, a bot-mitigation block, a partial content gate), never dropped to make the success rate look better.

  • A live re-audit caught a real fact change in the wild: GitHub's #1 trending COBOL repository swapped entirely (a 3,588-star repo to a 253-star one) between two live captures taken a day apart, disclosed in place rather than silently overwritten
  • 12 originally single-source rows deepened with a genuinely independent second method (Wikidata, live REST API headers, raw CPython/PEP GitHub source, NIH PubChem) — one traced a Wikidata citation to a real GitHub commit and found Rust's exact creation date, July 23, 2006, which appears nowhere in Wikipedia's own prose
  • New rows include a genuinely unplanned Wikipedia Special:Random navigation and a third real access pattern beyond clean success/failure: Instagram renders a real public follower count for an unauthenticated visitor but gates the actual content behind a forced sign-up modal

Fact-Verification Dataset

Real claims, two independent sources each, honest verdicts including disagreement

$369 once

Two hundred thirty-seven real-world claims across AI/tech, business, science, medicine, history, geography, culture, sports, entertainment, law, and government, each checked against a real cited source AND a second, genuinely independent source, with the two sources' agreement recorded explicitly. Built specifically to probe plausible-sounding falsehoods and implausible-sounding truths, not just easy trivia — and to surface the cases where two independent sources actually disagree, which a single-source lookup would never catch. A 2026-07-12 expansion added 37 new claims leaning into current events genuinely beyond any model's training cutoff (Nvidia's fiscal-2026 earnings, the 2026 Winter Olympics medal table, Super Bowl LX, the 98th Academy Awards), verified live via real API calls, Wikidata SPARQL queries, and independent web sources, alongside a full re-audit of the original 200 that found the existing data held up with zero new errors. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Expanded from 200 to 237 claims with 37 new dual-sourced rows on 2026 events verified live rather than recalled — Nvidia's $215.9B fiscal-2026 revenue, the 2026 Winter Olympics medal table, Super Bowl LX, and the 98th Academy Awards
  • A live re-audit of all 8 flagged source disagreements plus every AI/tech claim found the original 200 rows held up with zero new errors, while a live SEC EDGAR query caught Nvidia silently switching its revenue XBRL tag after fiscal 2022 — the same 'coverage gap, not staleness' pattern already documented for Wikidata
  • Model tested cold across all 237 claims: 214 correct, 16 confidently wrong (6.8%), and 7 honest hedges concentrated almost entirely in the new 2026 rows, where the model correctly declined to guess outcomes beyond its knowledge cutoff instead of fabricating an answer

Code-Reality Verification Dataset

Does this library API claim actually match what shipped — checked against the real published artifact

$319 once

Forty-two real claims about specific npm, PyPI, and (new this release) crates.io package APIs at exact versions, each checked by actually downloading and inspecting the real published tarball/sdist/crate file — and, for 38 of the 42 rows, confirmed a second, independent way by actually installing the real package and running the real code (or, for two Rust API-removal claims, capturing a real compiler error from trying to build against the claimed symbol). Targets the exact failure mode of hallucinated library APIs: functions, parameters, and behaviors that sound plausible but were never shipped, shipped differently, or existed only in a different version than the one claimed. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Every claim also tested live against Claude Sonnet 5 asked cold: 5 of 42 (12%) answers were confidently wrong on real npm/PyPI/crates.io package API questions
  • 38 of 42 rows independently confirmed two separate ways — real source inspection plus actually running the installed code, or for two Rust rows, a real compiler error — not just a text match
  • New crates.io ecosystem plus real dynamic execution caught a live nuance: on current Node.js, require()-ing an ESM-only package like chalk or node-fetch no longer throws outright — it silently returns a broken module object instead, catchable only by actually running it

Citation-Graph Verification Dataset

Does Paper A really cite Paper B, and is B legally quotable — checked end to end

$319 once

43 real multi-hop citation claims of the form "Paper A cites a paper titled X," each verified by walking Paper A's actual OpenAlex reference graph (not just checking that both papers individually exist) and, where the citation is real, checking Paper B's legally-quotable open-access status via Unpaywall. This pass re-audited every existing row live, added Semantic Scholar as a genuinely independent second citation-graph provider wherever it added real coverage, caught and openly corrected one real error from the original release (a citation OpenAlex's own graph was silently missing), and expanded into new sub-areas -- NLP embeddings, biomedical segmentation, distributed systems, clinical vaccine trials, classical neurophysiology, modern LLMs, diffusion models, and protein structure prediction -- including a live-discovered case where OpenAlex's own citation-graph data contains a chronologically impossible "citation," which a naive automated checker would wrongly certify as real.

  • A second independent citation-graph provider (Semantic Scholar) caught a real error the original single-method release missed: 'Attention Is All You Need' does genuinely cite the 1997 LSTM paper -- confirmed against the actual paper text -- even though OpenAlex's own reference graph silently omits it.
  • One new row documents a live OpenAlex data-integrity anomaly: GPT-3's real reference-graph record currently contains a chronologically impossible 'citation' to a paper dated six years after GPT-3 was published, which naive ID-and-title matching (including this dataset's own base pipeline) would wrongly certify as a real citation.
  • Open-access license confirmation keeps getting harder under live re-verification, not easier: the dataset's only prior 'confirmed CC-licensed' row didn't survive a live Unpaywall recheck (the license data had drifted since the original build) and zero of 43 rows now carry a confirmed reusable license.

UK Companies House Verification Dataset

Director, incorporation, status, and beneficial-ownership claims checked against the real UK company register

$319 once

55 real claims about UK-registered companies (mostly AI/tech firms) — director identity, incorporation dates, company status, registered addresses, beneficial ownership (PSC), cross-company appointment history, previous company names, SIC industry classification, filing history, and registered charges — each independently checked against the live UK Companies House register via both its free public website and its authenticated REST API. 82% of rows are backed by two genuinely independent, separately-coded verification paths; the rest are honestly single-sourced because no second method exists (PSC and appointment data have no public-HTML equivalent at all). Includes a real bug found and fixed in the dataset's own verification tooling this release, plus live model-graded testing on every row: asked cold with no tools, Claude Sonnet 5 hedges correctly on 95% of these claims and hallucinates a specific wrong answer on 2%, showing why a live register lookup beats an LLM's memory for this kind of fact.

  • Famous founders are repeatedly absent from their own UK subsidiary's board: Dario Amodei isn't a director of Anthropic Limited, Elon Musk isn't a director of XAI UK Limited, and Michael Truell (Anysphere/Cursor) isn't a director of Anysphere UK Ltd — the same real pattern confirmed a third time this release.
  • Coincidental company-name collisions are a systematic, repeatable trap, not a one-off: XAI UK Limited (2012) and Groq UK Limited (2004) both predate the AI companies they sound like by a decade or more, and 'Signal AI Ltd' (incorporated Nov 2025) is a completely unrelated company — the real Signal AI is legally registered as 'Signal Media Limited.'
  • This release found and fixed a real bug in its own verification code: a name-matching helper required exact-token equality, so ordinary claims like 'Vishal Marria' silently failed to match Companies House's fuller registered name 'MARRIA, Vishal Kumar' — caught, root-caused, fixed, and re-verified live, not papered over.

AI Safety / Refusal Testing Dataset

Cross-model-tier refusal divergence and consistency, not a single one-shot test

$299 once

Thirty real prompts drawn from two published, MIT-licensed AI-safety benchmarks (JailbreakBench, HarmBench), each individually screened against a strict category allowlist (several on-allowlist-labeled candidates, including a nuclear-weapon prompt mislabeled "Government decision-making," were excluded on individual review) and tested fresh against three model capability tiers with repeated sampling per tier (180 total live invocations), not a single one-shot test against one model. Every row now also carries a second, independent verification method — a live re-fetch of the source benchmark's own raw data confirming an exact prompt/category match — after a full re-audit of the original 21 rows found zero drift. Strictly scoped to lower-severity categories only — nothing touching weapons, hacking, violence, sexual content, trafficking, or self-harm — and stores only classifications, never generated harmful content.

  • Every row now carries two independent verification methods: a live 3-tier x 2-sample behavioral test (180 cells) plus a live source-fidelity re-fetch confirming the prompt/category exactly matches the original benchmark's own data file
  • Individual per-prompt screening actively excludes on-allowlist-labeled prompts that fail deeper review — this pass caught and dropped a nuclear-weapon-construction prompt filed under "Government decision-making," an organ-trafficking prompt, and three others, rather than trusting the source benchmark's own category labels
  • Still only one tier-divergent prompt in 30: an investment-advice bot request where haiku refuses, sonnet reframes into a disclaimer-heavy tool, and opus (with agentic tool access) locates and ships a working artifact — capability going up, caution going down, for this one category only

Cross-Source Verification Dataset

Two independent real sources, checked against each other, not just one

$319 once

Wikidata vs. live web; SEC filings vs. marketing claims; and, new in this expansion, one acquirer's audited SEC filing vs. another company's widely-reported deal-announcement figure for the same transaction. 42 real cases checking whether a structured official source actually agrees with a second, independent real source — and, where possible, a third. This pass re-verified all 27 original rows live with zero errors found, added a genuine second SEC-filed source to two rows that originally hit Wikidata's total silence (surfacing real gaps between announced and audited deal sizes), and expanded into 15 new rows covering newer AI companies and real corporate M&A, including a materially wrong headline acquisition-price figure caught by comparing an acquirer's own audited books against press coverage.

  • 42 rows across 3 methods (Wikidata-vs-live, marketing-vs-SEC-filing, and a new acquirer-filing-vs-deal-headline cross-check); the 2026-07-14 adversarial audit re-derived all 42 rows live, found every structured SEC/Wikidata value exact, and fixed 10 claimed-side citation errors in place with disclosed correction notes
  • New acquirer's-own-10-K vs press-release deal-price checks caught a real, material $1.03B (16%) gap for IBM's HashiCorp acquisition ($7.433B audited total consideration vs. the widely-repeated $6.4B headline) and resolved two prior Wikidata coverage gaps (Meta/Scale AI, CoreWeave/Weights & Biases) by finding the SEC-filed figures diverge from press-reported deal sizes
  • 7 new company rows (Perplexity, Groq, Character.AI, Cursor/Anysphere, Harvey AI, etc.) found Wikidata's ownership/CEO coverage gaps concentrate hardest on newer, fast-moving AI startups — 5 of 7 have zero structured statement at all, not just a stale one, including a live, unclosed $60B SpaceX-Cursor acquisition Wikidata hasn't recorded either way

Historical Event/Date/Figure Facts Verification

Historical event, date, and attribution claims independently checked against Wikidata's live structured statements, not model memory

$289 once

Forty-two real historical claims — a battle date, an assassination's causal framing, who really invented the telephone — each checked against Wikidata's real, live, CC0-licensed structured statements for the specific entity in question, including a purpose-built technique for compound multi-date events (space missions, voyages) where a naive property lookup would silently grab the wrong sub-event's date. most claims match, a documented minority diverge (off-by-one-day/year errors, invention misattributions, an overstated single-cause claim), and 1 is honestly marked unverifiable. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Every claim also tested live against Claude Sonnet 5 asked cold: 0 of 34 answers were wrong outright, the strongest clean-recall performance found anywhere in this catalog
  • Invention-attribution folklore fails twice, independently: Wikidata lists Reis, Gray, Bell, and Meucci for the telephone (no Edison at all), and five credited inventors for the incandescent light bulb
  • Three genuine "almost right" date errors caught by checking the exact claimed day, not just the year — the Titanic struck the iceberg April 14 but sank April 15; Lincoln was shot April 14 but died April 15
  • Wikidata's own "has cause" statement for World War I resolves to an item literally labeled "multiple causes," directly contradicting the popular single-trigger claim that Franz Ferdinand's assassination alone caused the war

English Language & Literary Attribution Verification

Famous quotes and book-publication claims checked against real public-domain full text and live Wikidata, not recalled

$289 once

Forty-two claims, each traced back to a real, live-fetched source: exact substring search against the real full plain-text of 25 public-domain books/plays via Project Gutenberg for quote claims, and live Wikidata publication-date/authorship checks. 16 of the original 35 rows (46%) are genuine divergences — misquotes, mixed-up authors, and swapped years that an LLM's training-data recall would very plausibly get wrong with high confidence — a count that grew by one when this dataset's own adversarial audit held a row to its own punctuation-strict standard and flipped it. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Every claim also tested live against Claude Sonnet 5 asked cold: 5 of 35 (14%) answers were confidently wrong, on top of the 16 divergences already found in the underlying source claims
  • "Elementary, my dear Watson" does not appear anywhere in Arthur Conan Doyle's original canon — confirmed by full-text-searching all 9 real Doyle books on Gutenberg, not just one
  • Oscar Wilde's "I can resist everything except temptation" is real but misattributed: not in The Picture of Dorian Gray at all, but verbatim in his play Lady Windermere's Fan
  • A live Wikidata check catches two planted plausible-sounding errors cleanly: The Great Gatsby claimed as the same year as Ulysses (1922; real value 1925), and Wuthering Heights claimed as Charlotte Brontë's (real author: her sister Emily)

Geographic Facts Verification

Country capitals, populations, land areas, borders, and physical-geography superlatives checked against two independent live sources

$289 once

Forty real geographic claims — capitals, population figures with year, land areas, bordering countries, and physical-geography superlatives — each cross-checked against two live, independently maintained sources: the CIA World Factbook's last public snapshot (public domain, US government work) and live Wikidata. Surfaces genuinely current, non-obvious findings, including a stale Wikidata population statement that would get "which country is more populous" wrong, and a currently unsettled naming dispute over North America's highest peak. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Every claim also tested live against Claude Sonnet 5 asked cold: 2 of 30 (7%) answers were confidently wrong
  • Wikidata's population statement for India is stuck at a 2020 figure while China's is current to 2025 — a tool trusting Wikidata alone would wrongly show China still ahead, when fresher 2025 estimates show India has actually overtaken China
  • Wikidata's entity for North America's highest point currently has no English label at all, reflecting a live, unsettled dispute after a 2025 US federal order reverted the peak's official name
  • The CIA World Factbook's official site was itself taken offline by the CIA in Feb 2026 — confirmed live via redirect to a "farewell" notice — forcing this product to source from the last public snapshot via a long-running open mirror

Language & Runtime Semantics Verification

Claims about Python/JavaScript runtime behavior, each settled by actually running the code or fetching the official changelog

$309 once

Forty claims about sort stability, integer/float precision, division/modulo sign semantics, string encoding edge cases, and "which version introduced feature X," verified by literally executing python3/node and capturing the real output, or by live-fetching the language's own official changelog. 31 of 40 rows carry real dual-interpreter executed code with literal output; the rest carry a live-fetched changelog snippet. Includes 5 deliberately adversarial rows built from real, commonly-stated wrong claims, to prove the method actually catches errors. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Every claim also tested live against Claude Sonnet 5 asked cold: 2 of 40 (5%) answers were confidently wrong on live-executed runtime-behavior questions
  • Node.js's global fetch() shipped in 18.0.0 explicitly marked "experimental" and wasn't promoted to stable until 21.0.0 — a real 3-major-version gap a compressed claim like "Node 18 shipped stable fetch" gets checkably wrong
  • Running 5 % -3 in both live interpreters shows Python's % takes the sign of the divisor while JavaScript's % takes the sign of the dividend — same operands, silently different answers
  • Python's `nan in [nan]` returns True even though `nan == nan` is False, because CPython's list-membership test tries object identity before falling back to equality — isolated by testing against a second, genuinely different nan object

SEC EDGAR Financial-Filing Verification

Real claims about public companies' revenue, net income, EPS, and filing dates checked live against SEC EDGAR's structured XBRL data

$319 once

Forty-one real claims about specific public companies' reported financials — revenue, net income, diluted EPS, shares outstanding, R&D spend, and 10-K filing dates — each checked against a live call to SEC EDGAR's XBRL "company facts" API. Covers 19 major companies across tech, retail, finance, energy, and pharma, including already-filed results more current than most model training cutoffs, a real cross-filing discrepancy between a company's own 10-K and proxy statement, and a currently-live SEC ticker-to-CIK mapping bug affecting a Dow 30 company. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Every claim also tested live against Claude Sonnet 5 asked cold: 3 of 41 (7%) answers were confidently wrong, and 8 more were honestly hedged rather than guessing an exact dollar figure
  • SEC's own live ticker-to-CIK map currently resolves ticker XOM to a newly-registered holdco with zero XBRL data, while decades of real Exxon financials remain under the original CIK — confirmed live and documented as paired unverifiable/match rows
  • Netflix's original FY2024 10-K tagged diluted EPS at $19.83; after Netflix's real 10-for-1 stock split, its FY2025 10-K retroactively restates the same FY2024 comparative to $1.98 — a genuine ~90% difference for the same underlying earnings
  • JPMorgan Chase's own audited 10-K tags net income at exactly $58,471,000,000, while its own later proxy statement tags the same nominal figure at a rounded $58,500,000,000 — a real $29 million gap between two of the company's own SEC filings

FDA Drug/Device Safety Verification

Real recall, adverse-event, and boxed-warning claims checked live against openFDA, not against memory

$319 once

Forty real-world drug and medical-device safety claims — recalls, FDA Adverse Event Reporting System/MAUDE report counts, and current FDA-mandated boxed-warning/label text — each independently checked against openFDA, the FDA's own live structured-data API, across four distinct endpoints. Deliberately distinct from this catalog's Clinical Trial Outcomes Verification product, which covers trial results and approvals, not real-world safety data. Built for health-AI products where a wrong safety answer is a high-stakes hallucination, not trivia. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Every claim also tested live against Claude Sonnet 5 asked cold: 4 of 28 (14%) answers were confidently wrong on real drug-safety questions, exactly the kind of hallucination a health-AI product can't afford
  • Two structurally-identical-looking openFDA queries for the same drug/reaction pair returned different adverse-event counts (8,655 vs 8,652) — a real, reproducible discrepancy between two common query shapes
  • Pulling a drug's full FDA label version history surfaced a genuine anomaly: a boxed warning removed by FDA action in 2016 briefly reappears in one intermediate label version, sandwiched between two versions that don't have it
  • A drug's current FDA label and independent adverse-event reporting corroborate the same safety signal from two completely different sources — the same reaction is the #1 most-reported FAERS term for that drug

Looking for the full company landscape dataset instead?

171 companies across 14 categories, sourced and cited.

See landscape dataset pricing