Browse the packaged proof evaluation and the broader catalog. Buy the 500-case drug proof dataset below, or review samples and request a scoped package for another domain.
Packaged dataset
A historical MED-RT source-grounding evaluation: 500 unique questions, verified reference answers, source provenance, and recorded model grades.
This is separate from the retired 338-row diagnostic export. It is not clinical advice, an independently blinded holdout, or evidence of model improvement.
Browse the catalog
Explore the existing inventory by task. The 500-case drug proof package above has automatic delivery. Other datasets require a reviewed scope and delivery agreement; larger remediation and validation assets remain pending publication.
Reference counts describe the listed export, not the combined size of every component in a product family. Component counts, source dates and limitations are in the dataset details.
Full catalog and component pricing →Prices below are reference prices for scoped packages unless an offer links directly to checkout. They do not imply automatic delivery. Larger locally validated assets remain pending publication; packaging, licensing and delivery must be agreed first.
Fresh proof evaluations that measure unsupported or incorrect answers.
Independent source-backed sets to test whether a weakness is systematic.
Separate training data built from confirmed failure modes.
Untouched examples for post-training improvement and regression checks.
Verified assets for a customer-specific domain or hallucination mode.
Flagship examples of the hallucination audit, remediation, and held-out validation framework.
Proof, remediation, DPO, and held-out data for PubMed metadata grounding and calibrated abstention
500 actual rows
500-row PubMed-backed proof evaluation with source-aware deterministic grading.
Request package9,573 actual rows
Validated remediation rows targeting observed refusal, metadata identity, and retraction/correction failures.
Request package970 actual rows
Untouched held-out rows excluded from proof and remediation PMIDs, DOIs, source records, and source URLs.
Request package11,043 actual rows
500-row proof eval, 9,573 remediation rows, and 970 held-out validation rows. No model improvement result is included.
Request packageAny larger PubMed package requires separate source-supply and validation review.
Request packagesource_aware_graded
Source-aware result: 161 correct, 34 substantive incorrect, and 305 refusal or uncertainty responses.
validated
9,145 unique PMIDs, 9,122 unique DOIs, and 2,467 unique journals.
validated
970 unique PMIDs, 920 unique DOIs, and 553 unique journals.
proof_derived_observed_failures
305 refusal-derived pairs and 34 substantive wrong-answer-derived pairs. No fabricated rejected responses.
public_preview
Public-safe preview. It does not include remediation or held-out datasets in full.
A PubMed-backed product family for citation metadata grounding and refusal calibration. On a 500-row PubMed-backed proof evaluation, the tested model achieved 32.2% accuracy. The source-aware result found 305 unnecessary refusal or uncertainty responses on answerable PubMed metadata and 34 substantive incorrect responses. The package includes 9,573 validated remediation rows, 970 untouched held-out rows, and 339 real observed proof-failure DPO pairs. Built to help models answer source-supported PubMed metadata while preserving appropriate abstention when metadata is genuinely unavailable. No model-improvement claim is made yet.
OpenAlex-backed proof, remediation, and held-out data for scholarly metadata grounding
43 actual rows
Existing published JSON product with verified citation/source-grounding rows.
Request package1,000 actual rows
Fresh OpenAlex-backed proof set with completed chatgpt-web browser run and source-aware deterministic grading.
Request package10,000 actual rows
Validated metadata-grounding training records targeting observed OA-status, author-attribution, and DOI-identity failures.
Request package1,000 actual rows
Untouched OpenAlex metadata validation rows reserved from proof and remediation generation.
Request package12,000 actual rows
Proof eval, 10k remediation data, and held-out validation packaged together. No trained-model improvement result is included yet.
Request packageLarger citation metadata assets require a separate scale gate; no 25k citation artifact is claimed here.
Request packageavailable
Original verified citation/source-grounding export.
validated and model-graded locally
Source-aware grading: 869 correct, 131 incorrect, 0 refused, 0 ungraded.
validated locally
Targets observed OA-status, first-author, and DOI metadata errors; no proof prompt/work overlap.
validated locally; reserved from remediation
Untouched held-out rows with zero proof/remediation overlap in validation.
A scholarly-source grounding product family built from the OpenAlex works API. The legacy 43-row citation-graph export remains available as a small verified eval, but the commercial package now centers on a fresh 1,000-row OpenAlex-backed proof evaluation with completed browser-assisted chatgpt-web responses and source-aware deterministic grading. Overall accuracy was 86.9% (869/1,000), with a 13.1% error rate; the weakness is narrowly scoped to OpenAlex metadata grounding, not a blanket citation-graph failure. Citation relationship and direction tasks were generally strong, while OpenAlex open-access status was the strongest observed weakness at 31.33% accuracy. A separate 10,000-row remediation/training dataset, 1,000-row untouched held-out validation set, and 115 observed-failure DPO examples are available. No post-training improvement claim is made yet.
Proof, remediation, and held-out data for period-specific financial fact grounding from SEC EDGAR XBRL
41 actual rows
Existing published JSON export with verified SEC financial-filing rows.
Request package500 actual rows
Fresh SEC EDGAR proof set with completed chatgpt-web browser run and deterministic source-aware grading.
Request package10,000 actual rows
Validated SEC companyfacts/companyconcept remediation records targeting observed numeric, temporal, and calculation failures.
Request package1,000 actual rows
Untouched SEC EDGAR validation rows excluded from proof and remediation source records.
Request package11,500 actual rows
500-row proof eval, 10,000 remediation rows, and 1,000 held-out validation rows. No model improvement result is included yet.
Request packageRequires a separate source-supply and validation review before any larger package is claimed.
Request packagepublished
Historical verified SEC financial-filing export.
graded_confirmed_weakness
Fresh proof set with 27.2% chatgpt-web accuracy and deterministic SEC-source grading.
passed
Training-oriented records targeting observed numeric, temporal, and calculation failures.
passed
Untouched validation rows not used for remediation generation.
proof_derived_observed_failures
Rejected responses are real observed proof-run errors, not synthetic negatives.
public_preview
Proof-derived public sample. It does not include remediation or held-out rows.
A SEC EDGAR financial-grounding product family built from companyfacts, companyconcept, submissions metadata, and filing-index source URLs. The original 41-row diagnostic export remains available, and a fresh 500-row proof evaluation has now been completed with browser-assisted chatgpt-web responses and deterministic SEC-source grading. The proof run measured 136/500 correct, 364/500 incorrect, 27.2% accuracy, and 72.8% error, concentrated in period-specific numeric financial facts, date and period confusion, and YoY calculations. A separate 10,000-row SEC remediation/training dataset and 1,000-row untouched held-out validation set now exist. Built to help teams train against observed SEC financial hallucination modes and designed for before/after validation on untouched held-out data. No post-training improvement claim is made yet.
Real recall, adverse-event, boxed-warning, and label-section claims checked live against openFDA
40 actual rows
Existing published JSON product with verified drug/device safety rows.
Request package1,000 actual rows
Fresh openFDA label proof set with completed chatgpt-web browser run and strict deterministic grading caveat.
Request package10,000 actual rows
Validated label-section remediation records targeting observed indications_and_usage extraction failures.
Request package1,000 actual rows
Untouched openFDA label rows reserved from remediation generation.
Request package12,000 actual rows
Proof eval, 10k remediation data, and held-out validation packaged together. No post-training improvement result is included yet.
Request packageLarger openFDA label assets require a separate source-supply and validation gate.
Request packagepublished
Existing recall, FAERS/MAUDE, and boxed-warning checks.
model-graded
Full 1,000-row browser-assisted chatgpt-web run imported and graded.
validated
Training data targets complete indications_and_usage section grounding.
reserved
Untouched validation rows reserved for post-training evaluation.
Forty real-world drug and medical-device safety claims: recalls, FDA Adverse Event Reporting System/MAUDE report counts, and current FDA-mandated boxed-warning/label text: each independently checked against openFDA, the FDA's own live structured-data API, across four distinct endpoints. A separate 1,000-row openFDA label proof evaluation has now been run through the browser-assisted ChatGPT workflow. Under strict deterministic grading, chatgpt-web answered 844/1,000 correctly; the 156 non-correct rows were concentrated in hard indications_and_usage label-section extraction. A 10,000-row remediation/training artifact and 1,000-row untouched held-out validation set now exist for that observed label-section weakness. No post-training improvement claim is made yet.
Drug-indication source-grounding data: graded diagnostic set, proof eval, 10k remediation data, and held-out eval
338 actual rows
Existing published JSON product with model-graded diagnostic/eval rows.
Request package500 actual rows
Packaged 500 unique cases, provenance, historical grades, manifest and internal-use license. Automatic delivery after confirmed payment.
View checkout — $99910,000 actual rows
Validated training/remediation records with train/dev/test split and source-record leakage checks.
Request package1,000 actual rows
Untouched held-out RxNav evaluation set reserved from remediation generation.
Request package11,500 actual rows
Proof eval, 10k remediation data, and held-out validation packaged together. No post-training improvement result is included yet.
Request packageCustomer-specific diagnostic, proof, remediation, and held-out study scoped to the buyer model/domain.
Request packagehistorical archive; retired from sale
Fully model-graded; source-grounded against RxNav/RxClass MED-RT.
packaged for purchase with automatic delivery
Browser-assisted ChatGPT run: 108 correct, 392 incorrect/hallucinated, 0 refused, 0 ungraded.
validated locally; pending publication
Deterministic 9,000/500/500 train/dev/test split; factual answers derived from verified source rows.
validated locally; reserved from remediation
Zero RxCUI and row-ID overlap with the 10k remediation dataset.
not yet run
Infrastructure exists, but no drug-indication post-training result has been measured.
A source-grounded drug-indication product family built from the National Library of Medicine RxNav/RxClass MED-RT DISEASE may_treat API. The published 338-row diagnostic/evaluation set is fully model-graded and found a measured 33.7% hallucination rate (114/338) when asked cold with no source access. A fresh 500-row proof evaluation was then run through the browser-assisted ChatGPT workflow and deterministically graded against RxNav/MED-RT, finding 392/500 incorrect or hallucinated answers. A separate 10,000-record remediation/training dataset and a 1,000-row untouched held-out evaluation now exist with source-record leakage checks. No post-training improvement claim is made yet; the held-out measurement step still needs an actual trained model run.
Smaller verified datasets for evaluation and regression testing. These are not presented as full remediation systems unless proof/remediation assets exist.
Theorem, constant, and attribution claims checked by real computation, not recall
Small verified evaluation export; no separate proof/remediation asset is packaged for this product yet.
Forty-eight real mathematical claims (famous theorem statements, numeric sequence/constant properties, computational results, open-problem status checks, and "who first proved X" attributions), each checked against code actually written and run (Python sympy/mpmath, a from-scratch BigInt library, and an exhaustive brute-force combinatorial search), the real Wikidata knowledge graph, real published papers via OpenAlex, and, for one live current-events check, primary web sources fetched live. An AI model asked to check its own math claim just restates it more confidently; this dataset actually computes, queries, and looks the answer up, and says so honestly on the rare row where a second independent check genuinely doesn't exist. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.
Physical constants and discovery claims checked against NIST, OpenAlex, Wikidata, and real independent computation
Small verified scientific-factuality evaluation export.
Forty-three real claims about physical constants, first-discovery/first-measurement attributions, and specific published experimental results across physics, chemistry, astronomy, biology, and earth science, each checked against NIST's live CODATA fundamental-constants database, the actual paper reporting a result via OpenAlex, Wikidata's structured discoverer/date/value statements, or, new this pass, real independently executed computation deriving one fundamental constant from others. Fills the gap between Groundtruth Data's clinical-trial and citation-graph products, which don't touch general physical-science claims, and is now the catalog's most methodologically diverse verification product. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.
Patent and IP claims checked against the live USPTO record, not what a company says about them
Small verified IP/factuality evaluation export.
Forty real claims companies and named inventors have made about patents ("we hold a patent on X," "patent number Y covers Z," "first to patent W"), each checked against the actual live record at USPTO's Patent Public Search system: does the patent exist, is it assigned to the claimed entity, do its real granted claims actually cover what's being claimed, and is it still valid today. Every row is now also cross-checked against a second, independent USPTO data feed: the live Patent Assignment and PTAB docket history tracking what actually happened to a patent after it was granted, not just its bibliographic snapshot: which caught patents that quietly lapsed for non-payment years after being cited in litigation, ownership chains that don't match a company's own marketing page, and three real errors in this dataset's own earlier release, corrected in place and disclosed rather than silently fixed. The original 23 rows also carry a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.
"Case X held Y" claims checked against the real opinion text, not a paraphrase
Small verified legal-citation evaluation export.
Forty-two real claims of the form "Case X held Y," each verified by fetching the actual opinion text of Case X from Harvard Law School's Caselaw Access Project (CC0 public domain, confirmed live) and checking whether the claimed holding, quote, vote count, or date genuinely appears in the primary source: now cross-checked against a second independent source, Wikidata's structured case-law statements, on most rows. Grew from a 20-case constitutional-law pilot into a wider set spanning antitrust, torts, foundational property law, copyright, trademark, and federal appellate law, including this dataset's first non-Supreme-Court cases (New York's Court of Appeals, an 1805 property-law ruling, and the D.C. Circuit's Microsoft antitrust decision). CourtListener was deliberately not used, its commercial-resale licensing remains unconfirmed with Free Law Project. The original 20 rows also carry a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.
Does the trial data behind a drug claim actually back it up: checked against the real registry
Legacy diagnostic export is fixed-price. Expanded proof, remediation, held-out, and DPO package should be quoted as a clinical refusal-calibration system.
published
Existing clinical-source verification product.
source_aware_graded
Source-aware result: 45 correct, 4 substantive incorrect, and 451 unnecessary refusal or uncertainty responses.
validated
Built to teach supported answers when ClinicalTrials.gov contains the fact and calibrated abstention when the source is insufficient.
validated
Untouched held-out validation with zero proof or remediation NCT overlap.
Forty-seven real claims about clinical trial outcomes and FDA drug approvals remain available as the legacy diagnostic export. The expanded ClinicalTrials.gov Fact Grounding and Refusal Calibration package adds a 500-row proof evaluation, a 10,000-row remediation set, a 1,000-row untouched held-out set, and 455 observed-failure DPO pairs. The source-aware proof result showed 9.0% accuracy with 451 unnecessary refusals or uncertainty responses on directly verifiable ClinicalTrials.gov facts and 4 substantive wrong answers. This is framed as refusal calibration and registry fact grounding, not as a fabricated-fact hallucination benchmark.
Federal contracting claims checked against the actual USASpending.gov award record
Small verified government-contract evaluation export.
Forty real claims companies make about their US federal and defense contracting relationships: "prime DoD contractor," specific dollar figures, named agency relationships: each checked against the real award data on USASpending.gov, the federal government's own system of record for who it has actually paid, with a second independent method (real SEC EDGAR filings) added wherever one genuinely exists. Surfaces real, checkable failure modes: contract-ceiling figures presented as if already earned, a real NYSE-listed company with billions in current federal work that's entirely invisible under its own name because every dollar is still filed under its pre-merger legal entity, and: new this pass: a household-name defense prime whose federal business is invisible to a simple name search because the identical legal entity is registered under more than 100 separate, un-consolidated USASpending IDs. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.
Real claims, two independent sources each, honest verdicts including disagreement
Broad verified factuality evaluation export with 237 rows.
Two hundred thirty-seven real-world claims across AI/tech, business, science, medicine, history, geography, culture, sports, entertainment, law, and government, each checked against a real cited source AND a second, genuinely independent source, with the two sources' agreement recorded explicitly. Built specifically to probe plausible-sounding falsehoods and implausible-sounding truths, not just easy trivia: and to surface the cases where two independent sources actually disagree, which a single-source lookup would never catch. A 2026-07-12 expansion added 37 new claims leaning into current events genuinely beyond any model's training cutoff (Nvidia's fiscal-2026 earnings, the 2026 Winter Olympics medal table, Super Bowl LX, the 98th Academy Awards), verified live via real API calls, Wikidata SPARQL queries, and independent web sources, alongside a full re-audit of the original 200 that found the existing data held up with zero new errors. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.
Does this library API claim actually match what shipped: checked against the real published artifact
Small verified code/package factuality evaluation export.
Forty-two real claims about specific npm, PyPI, and (new this release) crates.io package APIs at exact versions, each checked by actually downloading and inspecting the real published tarball/sdist/crate file: and, for 38 of the 42 rows, confirmed a second, independent way by actually installing the real package and running the real code (or, for two Rust API-removal claims, capturing a real compiler error from trying to build against the claimed symbol). Targets the exact failure mode of hallucinated library APIs: functions, parameters, and behaviors that sound plausible but were never shipped, shipped differently, or existed only in a different version than the one claimed. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.
Director, incorporation, status, and beneficial-ownership claims checked against the real UK company register
Small verified corporate-registry evaluation export with authenticated-source verification.
55 real claims about UK-registered companies (mostly AI/tech firms): director identity, incorporation dates, company status, registered addresses, beneficial ownership (PSC), cross-company appointment history, previous company names, SIC industry classification, filing history, and registered charges: each independently checked against the live UK Companies House register via both its free public website and its authenticated REST API. 82% of rows are backed by two genuinely independent, separately-coded verification paths; the rest are honestly single-sourced because no second method exists (PSC and appointment data have no public-HTML equivalent at all). Includes a real bug found and fixed in the dataset's own verification tooling this release, plus live model-graded testing on every row: asked cold with no tools, Claude Sonnet 5 hedges correctly on 95% of these claims and hallucinates a specific wrong answer on 2%, showing why a live register lookup beats an LLM's memory for this kind of fact.
Cross-model-tier refusal divergence and consistency, not a single one-shot test
Pilot-scale verified refusal-testing export.
Thirty real prompts drawn from two published, MIT-licensed AI-safety benchmarks (JailbreakBench, HarmBench), each individually screened against a strict category allowlist (several on-allowlist-labeled candidates, including a nuclear-weapon prompt mislabeled "Government decision-making," were excluded on individual review) and tested fresh against three model capability tiers with repeated sampling per tier (180 total live invocations), not a single one-shot test against one model. Every row now also carries a second, independent verification method: a live re-fetch of the source benchmark's own raw data confirming an exact prompt/category match: after a full re-audit of the original 21 rows found zero drift. Strictly scoped to lower-severity categories only: nothing touching weapons, hacking, violence, sexual content, trafficking, or self-harm: and stores only classifications, never generated harmful content.
Two independent real sources, checked against each other, not just one
Small verified cross-source consistency evaluation export.
Wikidata vs. live web; SEC filings vs. marketing claims; and, new in this expansion, one acquirer's audited SEC filing vs. another company's widely-reported deal-announcement figure for the same transaction. 42 real cases checking whether a structured official source actually agrees with a second, independent real source: and, where possible, a third. This pass re-verified all 27 original rows live with zero errors found, added a genuine second SEC-filed source to two rows that originally hit Wikidata's total silence (surfacing real gaps between announced and audited deal sizes), and expanded into 15 new rows covering newer AI companies and real corporate M&A, including a materially wrong headline acquisition-price figure caught by comparing an acquirer's own audited books against press coverage.
Historical event, date, and attribution claims independently checked against Wikidata's live structured statements, not model memory
Small verified history factuality evaluation export.
Forty-two real historical claims: a battle date, an assassination's causal framing, who really invented the telephone: each checked against Wikidata's real, live, CC0-licensed structured statements for the specific entity in question, including a purpose-built technique for compound multi-date events (space missions, voyages) where a naive property lookup would silently grab the wrong sub-event's date. most claims match, a documented minority diverge (off-by-one-day/year errors, invention misattributions, an overstated single-cause claim), and 1 is honestly marked unverifiable. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.
Famous quotes and book-publication claims checked against real public-domain full text and live Wikidata, not recalled
Small verified literary attribution evaluation export.
Forty-two claims, each traced back to a real, live-fetched source: exact substring search against the real full plain-text of 25 public-domain books/plays via Project Gutenberg for quote claims, and live Wikidata publication-date/authorship checks. 16 of the original 35 rows (46%) are genuine divergences: misquotes, mixed-up authors, and swapped years that an LLM's training-data recall would very plausibly get wrong with high confidence: a count that grew by one when this dataset's own adversarial audit held a row to its own punctuation-strict standard and flipped it. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.
Country capitals, populations, land areas, borders, and physical-geography superlatives checked against two independent live sources
Small verified geography factuality evaluation export.
Forty real geographic claims: capitals, population figures with year, land areas, bordering countries, and physical-geography superlatives: each cross-checked against two live, independently maintained sources: the CIA World Factbook's last public snapshot (public domain, US government work) and live Wikidata. Surfaces genuinely current, non-obvious findings, including a stale Wikidata population statement that would get "which country is more populous" wrong, and a currently unsettled naming dispute over North America's highest peak. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.
Claims about Python/JavaScript runtime behavior, each settled by actually running the code or fetching the official changelog
Small verified runtime-semantics evaluation export with an improvement-report scaffold.
Forty-six claims about sort stability, integer/float precision, division/modulo sign semantics, string encoding edge cases, mutability/identity gotchas, and "which version introduced feature X," verified by literally executing python3/node and capturing the real output, or by live-fetching the language's own official changelog. Most rows carry real dual-interpreter executed code with literal output; the rest carry a live-fetched changelog snippet. Includes several deliberately adversarial rows built from real, commonly-stated wrong claims, to prove the method actually catches errors. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.
Start with a suspected weakness and target capability; Groundtruth Data scopes diagnostic, proof, remediation, and held-out assets around it.
171 companies across 14 categories, sourced and cited.