Products

All datasets

Browse the packaged proof evaluation and the broader catalog. Buy the 500-case drug proof dataset below, or review samples and request a scoped package for another domain.

Packaged dataset

500-case Drug Indication Proof Evaluation

A historical MED-RT source-grounding evaluation: 500 unique questions, verified reference answers, source provenance, and recorded model grades.

  • JSONL evaluation and provenance files
  • Historical model responses and grades
  • SHA-256 manifest and internal-use license

This is separate from the retired 338-row diagnostic export. It is not clinical advice, an independently blinded holdout, or evidence of model improvement.

Browse the catalog

21 dataset and product-family entries.

Explore the existing inventory by task. The 500-case drug proof package above has automatic delivery. Other datasets require a reviewed scope and delivery agreement; larger remediation and validation assets remain pending publication.

Dataset / product familyListed reference rowsPurchase pathExplore
Mathematical Claims Verification48Request scope and delivery quoteView sample
Scientific Claims Verification43Request scope and delivery quoteView sample
Patent & IP Claims Verification40Request scope and delivery quoteView sample
Legal Citation & Case Law Verification42Request scope and delivery quoteView sample
Clinical Trial Outcomes Verification47Request scope and delivery quoteView sample
Government Contract Verification40Request scope and delivery quoteView sample
PubMed Citation Metadata and Refusal Calibration9,573Request scope and delivery quoteView sample
Web-Navigation Trajectories42Request scope and delivery quoteView sample
Fact-Verification Dataset237Request scope and delivery quoteView sample
Code-Reality Verification Dataset42Request scope and delivery quoteView sample
Citation / OpenAlex Metadata Grounding43Request scope and delivery quoteView sample
UK Companies House Verification Dataset55Request scope and delivery quoteView sample
AI Safety / Refusal Testing Dataset30Request scope and delivery quoteView sample
Cross-Source Verification Dataset42Request scope and delivery quoteView sample
Historical Event/Date/Figure Facts Verification42Request scope and delivery quoteView sample
English Language & Literary Attribution Verification42Request scope and delivery quoteView sample
Geographic Facts Verification40Request scope and delivery quoteView sample
Language & Runtime Semantics Verification46Request scope and delivery quoteView sample
SEC / EDGAR Financial Grounding41Request scope and delivery quoteView sample
FDA Drug/Device Safety Verification40Request scope and delivery quoteView sample
Drug Indication Diagnostic Archive (retired)338Historical diagnostic retired; 500-case package aboveView sample

Reference counts describe the listed export, not the combined size of every component in a product family. Component counts, source dates and limitations are in the dataset details.

Full catalog and component pricing →

Dataset details and reference pricing

Prices below are reference prices for scoped packages unless an offer links directly to checkout. They do not imply automatic delivery. Larger locally validated assets remain pending publication; packaging, licensing and delivery must be agreed first.

Hallucination Audit

Fresh proof evaluations that measure unsupported or incorrect answers.

Proof evaluations

Independent source-backed sets to test whether a weakness is systematic.

Hallucination Remediation

Separate training data built from confirmed failure modes.

Held-out evaluations

Untouched examples for post-training improvement and regression checks.

Custom study

Verified assets for a customer-specific domain or hallucination mode.

Core domain systems

Flagship examples of the hallucination audit, remediation, and held-out validation framework.

PubMed Citation Metadata and Refusal Calibration

Proof, remediation, DPO, and held-out data for PubMed metadata grounding and calibrated abstention

from $6,900 reference price
CapabilityPubMed citation metadata grounding and refusal calibration
Designationmixed
Version2026-08-31
Updated2026-09-01
SourcesNCBI PubMed E-utilities
FormatsJSONL records, JSONL eval, JSONL SFT, JSONL observed-failure DPO, held-out eval JSONL
Product typeproduct family
Pricing modelcomponent pricing
500Proof rowsPubMed-backed proof evaluation.
32.2%Proof accuracy161 correct out of 500 source-aware graded rows.
67.8%Non-correct305 refusal or uncertainty responses and 34 substantive incorrect responses.
9,573Remediation rowsValidated PubMed metadata and refusal-calibration records.
970Held-out rowsUntouched validation rows with zero proof or remediation overlap.
339DPO pairsRejected answers are real observed proof-run failures.
PubMed Metadata Proof Evalproof eval
$1,299 reference price

500 actual rows

500-row PubMed-backed proof evaluation with source-aware deterministic grading.

Request package
PubMed Metadata Remediationremediation training
$6,900 reference price

9,573 actual rows

Validated remediation rows targeting observed refusal, metadata identity, and retraction/correction failures.

Request package
PubMed Held-Out Validationheld out validation
$1,500 reference price

970 actual rows

Untouched held-out rows excluded from proof and remediation PMIDs, DOIs, source records, and source URLs.

Request package
PubMed Metadata Calibration Packfull improvement pack
$8,500 reference price

11,043 actual rows

500-row proof eval, 9,573 remediation rows, and 970 held-out validation rows. No model improvement result is included.

Request package
PubMed expansioncustom dataset
Custom quote

Any larger PubMed package requires separate source-supply and validation review.

Request package
PubMed proof evaluationproof eval
500 rows

source_aware_graded

Source-aware result: 161 correct, 34 substantive incorrect, and 305 refusal or uncertainty responses.

PubMed remediation recordsremediation
9,573 rows

validated

9,145 unique PMIDs, 9,122 unique DOIs, and 2,467 unique journals.

PubMed held-out validationheld out
970 rows

validated

970 unique PMIDs, 920 unique DOIs, and 553 unique journals.

PubMed observed-failure DPO poolremediation
339 rows

proof_derived_observed_failures

305 refusal-derived pairs and 34 substantive wrong-answer-derived pairs. No fabricated rejected responses.

Public PubMed samplesample
75 rows

public_preview

Public-safe preview. It does not include remediation or held-out datasets in full.

A PubMed-backed product family for citation metadata grounding and refusal calibration. On a 500-row PubMed-backed proof evaluation, the tested model achieved 32.2% accuracy. The source-aware result found 305 unnecessary refusal or uncertainty responses on answerable PubMed metadata and 34 substantive incorrect responses. The package includes 9,573 validated remediation rows, 970 untouched held-out rows, and 339 real observed proof-failure DPO pairs. Built to help models answer source-supported PubMed metadata while preserving appropriate abstention when metadata is genuinely unavailable. No model-improvement claim is made yet.

  • On a 500-row PubMed-backed proof evaluation, the tested model achieved 32.2% accuracy under source-aware deterministic grading
  • 305 of 339 non-correct responses were unnecessary refusal or uncertainty on answerable PubMed metadata
  • The proof also identified 29 substantive errors across 40 retraction/correction classification questions
  • The strongest refusal patterns were PMID identity, DOI recall, paper title, journal, publication date, first author, and author-list questions
  • The remediation package contains 9,573 validated rows spanning 9,145 unique PMIDs, 9,122 unique DOIs, and 2,467 unique journals
  • The held-out package contains 970 untouched rows spanning 970 unique PMIDs, 920 unique DOIs, and 553 unique journals
  • No claim is made that training has improved a model; held-out validation is reserved for a separate before and after experiment

Citation / OpenAlex Metadata Grounding

OpenAlex-backed proof, remediation, and held-out data for scholarly metadata grounding

from $299 reference price
CapabilityOpenAlex metadata grounding
Designationmixed
Version2026-08-30-proof-backed-10k
Updated2026-08-30
SourcesOpenAlex works API
FormatsJSON, JSONL proof eval, JSONL remediation records, JSONL SFT, JSONL DPO, JSONL held-out eval
Product typeproduct family
Pricing modelcomponent pricing
1,000OpenAlex proof evalFresh source-backed proof rows; 1,000 completed chatgpt-web responses
86.9%Source-aware accuracy869/1,000 correct under deterministic OpenAlex-aware grading
13.1%Source-aware error rate131/1,000 incorrect; Wilson 95% CI 11.15%-15.33%
31.33%OA-status accuracy47/150 correct; strongest observed failure family
10,000Remediation data7,023 unique works, 7,020 unique identifiers, validation passed
1,000Held-out evalUntouched OpenAlex metadata rows; overlap/leakage checks passed
Legacy citation-graph exportdiagnostic eval
$299 reference price

43 actual rows

Existing published JSON product with verified citation/source-grounding rows.

Request package
Citation / OpenAlex Metadata Proof Evalproof eval
$1,299 reference price

1,000 actual rows

Fresh OpenAlex-backed proof set with completed chatgpt-web browser run and source-aware deterministic grading.

Request package
OpenAlex Metadata Remediation 10kremediation training
$6,500 reference price

10,000 actual rows

Validated metadata-grounding training records targeting observed OA-status, author-attribution, and DOI-identity failures.

Request package
Citation Held-Out Validationheld out validation
$1,250 reference price

1,000 actual rows

Untouched OpenAlex metadata validation rows reserved from proof and remediation generation.

Request package
Citation Metadata Improvement Packfull improvement pack
$8,250 reference price

12,000 actual rows

Proof eval, 10k remediation data, and held-out validation packaged together. No trained-model improvement result is included yet.

Request package
Citation Metadata 25k+custom dataset
Custom quote

Larger citation metadata assets require a separate scale gate; no 25k citation artifact is claimed here.

Request package
Legacy citation-graph evaldiagnostic
43 rows

available

Original verified citation/source-grounding export.

OpenAlex metadata proof evaluationproof eval
1,000 rows

validated and model-graded locally

Source-aware grading: 869 correct, 131 incorrect, 0 refused, 0 ungraded.

OpenAlex metadata remediationremediation
10,000 rows

validated locally

Targets observed OA-status, first-author, and DOI metadata errors; no proof prompt/work overlap.

Citation metadata held-out validationheld out
1,000 rows

validated locally; reserved from remediation

Untouched held-out rows with zero proof/remediation overlap in validation.

A scholarly-source grounding product family built from the OpenAlex works API. The legacy 43-row citation-graph export remains available as a small verified eval, but the commercial package now centers on a fresh 1,000-row OpenAlex-backed proof evaluation with completed browser-assisted chatgpt-web responses and source-aware deterministic grading. Overall accuracy was 86.9% (869/1,000), with a 13.1% error rate; the weakness is narrowly scoped to OpenAlex metadata grounding, not a blanket citation-graph failure. Citation relationship and direction tasks were generally strong, while OpenAlex open-access status was the strongest observed weakness at 31.33% accuracy. A separate 10,000-row remediation/training dataset, 1,000-row untouched held-out validation set, and 115 observed-failure DPO examples are available. No post-training improvement claim is made yet.

  • The expanded 1,000-row proof run found a narrow, reproducible OpenAlex metadata-grounding weakness: 103 of 131 failures were open_access_status errors, while chronological false edges, nonreferences, and reverse-direction citation checks were all answered correctly.
  • The 10,000-row remediation dataset is built from fresh OpenAlex works, with 7,023 unique works, 7,020 unique identifiers, zero validation failures, zero proof prompt/work overlap, and leakage checks passed.
  • DPO data is limited to 115 genuine observed incorrect proof responses; rejected answers are not fabricated.
43 rows · JSON · remediation
Metadata and sample rowsDiscuss this dataset

SEC / EDGAR Financial Grounding

Proof, remediation, and held-out data for period-specific financial fact grounding from SEC EDGAR XBRL

from $299 reference price
CapabilitySEC financial fact grounding
Designationmixed
Version2026-08-30
Updated2026-08-30
SourcesSEC EDGAR companyfacts, SEC EDGAR companyconcept, SEC EDGAR submissions metadata, 10-K and 10-Q filing metadata
FormatsJSON, JSONL eval, JSONL remediation, JSONL SFT, JSONL observed-failure DPO
Product typeproduct family
Pricing modelcomponent pricing
500Proof rowsFresh SEC EDGAR proof evaluation.
27.2%Baseline accuracy136/500 correct on the completed chatgpt-web proof run.
72.8%Baseline error rate364/500 incorrect, 0 refusals.
68.7% to 76.5%Error 95% CIWilson interval from the 500-row proof result.
10,000Remediation rowsValidated SEC-source remediation/training records.
1,000Held-out rowsUntouched validation rows excluded from proof and remediation.
364Observed-failure DPORejected answers are real proof-run model errors.
Legacy SEC diagnostic exportdiagnostic eval
$299 reference price

41 actual rows

Existing published JSON export with verified SEC financial-filing rows.

Request package
SEC / EDGAR Financial Proof Evalproof eval
$1,299 reference price

500 actual rows

Fresh SEC EDGAR proof set with completed chatgpt-web browser run and deterministic source-aware grading.

Request package
SEC / EDGAR Remediation 10kremediation training
$6,900 reference price

10,000 actual rows

Validated SEC companyfacts/companyconcept remediation records targeting observed numeric, temporal, and calculation failures.

Request package
SEC / EDGAR Held-Out Validationheld out validation
$1,500 reference price

1,000 actual rows

Untouched SEC EDGAR validation rows excluded from proof and remediation source records.

Request package
SEC Financial Grounding Improvement Packfull improvement pack
$8,500 reference price

11,500 actual rows

500-row proof eval, 10,000 remediation rows, and 1,000 held-out validation rows. No model improvement result is included yet.

Request package
SEC 25k+ expansioncustom dataset
Custom quote

Requires a separate source-supply and validation review before any larger package is claimed.

Request package
Legacy SEC diagnostic exportdiagnostic
41 rows

published

Historical verified SEC financial-filing export.

SEC / EDGAR proof evalproof eval
500 rows

graded_confirmed_weakness

Fresh proof set with 27.2% chatgpt-web accuracy and deterministic SEC-source grading.

SEC / EDGAR remediation 10kremediation
10,000 rows

passed

Training-oriented records targeting observed numeric, temporal, and calculation failures.

SEC / EDGAR held-out validationheld out
1,000 rows

passed

Untouched validation rows not used for remediation generation.

SEC observed-failure DPO poolremediation
364 rows

proof_derived_observed_failures

Rejected responses are real observed proof-run errors, not synthetic negatives.

Public SEC samplesample
75 rows

public_preview

Proof-derived public sample. It does not include remediation or held-out rows.

A SEC EDGAR financial-grounding product family built from companyfacts, companyconcept, submissions metadata, and filing-index source URLs. The original 41-row diagnostic export remains available, and a fresh 500-row proof evaluation has now been completed with browser-assisted chatgpt-web responses and deterministic SEC-source grading. The proof run measured 136/500 correct, 364/500 incorrect, 27.2% accuracy, and 72.8% error, concentrated in period-specific numeric financial facts, date and period confusion, and YoY calculations. A separate 10,000-row SEC remediation/training dataset and 1,000-row untouched held-out validation set now exist. Built to help teams train against observed SEC financial hallucination modes and designed for before/after validation on untouched held-out data. No post-training improvement claim is made yet.

  • A fresh 500-row SEC proof evaluation found 136 correct and 364 incorrect chatgpt-web responses under deterministic SEC-source grading: 27.2% accuracy and 72.8% error
  • The strongest repeatable failures were period-specific numeric financial facts, fiscal dates, temporal comparisons, and YoY calculations; company identity and CIK/ticker tasks performed strongly and are kept as a small control group
  • Built to help teams train against observed SEC financial hallucination modes and designed for before/after validation on untouched held-out data
  • The 10,000-row remediation set spans 318 companies, 2,380 accessions, 16 XBRL concepts, and 1,461 fiscal periods, with zero proof or held-out accession overlap
  • The 1,000-row held-out validation set spans 119 companies, 447 accessions, 14 XBRL concepts, and 355 fiscal periods, and remains untouched for post-training measurement
  • SEC's own live ticker-to-CIK map currently resolves ticker XOM to a newly-registered holdco with zero XBRL data, while decades of real Exxon financials remain under the original CIK: confirmed live and documented as paired unverifiable/match rows
  • No claim is made that training has improved a model yet; held-out validation is reserved for that separate before/after experiment

FDA Drug/Device Safety Verification

Real recall, adverse-event, boxed-warning, and label-section claims checked live against openFDA

from $299 reference price
CapabilityFDA/openFDA safety and label factuality
Designationmixed
Version2026-08-29-proof-backed
Updated2026-08-29
SourcesopenFDA drug/label.json, openFDA drug/event.json, openFDA enforcement APIs, openFDA device APIs
FormatsJSON, JSONL, SFT JSONL for remediation
Product typeproduct family
Pricing modelcomponent pricing
1,000Label proof evalFresh source-verified openFDA drug label rows
84.4%Strict accuracy844/1,000 chatgpt-web browser-eval responses correct under deterministic grading
15.6%Strict error rate156/1,000 non-correct; Wilson 95% CI 13.5%-18.0%
10,000Remediation data9,000 train / 500 dev / 500 test; validation passed
1,000Held-out evalFresh openFDA label rows not used for remediation generation
Legacy FDA safety exportdiagnostic eval
$299 reference price

40 actual rows

Existing published JSON product with verified drug/device safety rows.

Request package
OpenFDA Label Proof Evalproof eval
$1,299 reference price

1,000 actual rows

Fresh openFDA label proof set with completed chatgpt-web browser run and strict deterministic grading caveat.

Request package
OpenFDA Label Remediation 10kremediation training
$5,900 reference price

10,000 actual rows

Validated label-section remediation records targeting observed indications_and_usage extraction failures.

Request package
OpenFDA Held-Out Validationheld out validation
$1,250 reference price

1,000 actual rows

Untouched openFDA label rows reserved from remediation generation.

Request package
OpenFDA Label Improvement Packfull improvement pack
$7,500 reference price

12,000 actual rows

Proof eval, 10k remediation data, and held-out validation packaged together. No post-training improvement result is included yet.

Request package
OpenFDA 25k+custom dataset
Custom quote

Larger openFDA label assets require a separate source-supply and validation gate.

Request package
Original safety diagnostic/evaldiagnostic
40 rows

published

Existing recall, FAERS/MAUDE, and boxed-warning checks.

OpenFDA label proof evaluationproof eval
1,000 rows

model-graded

Full 1,000-row browser-assisted chatgpt-web run imported and graded.

OpenFDA label-section remediationremediation
10,000 rows

validated

Training data targets complete indications_and_usage section grounding.

OpenFDA label-section held-out validationheld out
1,000 rows

reserved

Untouched validation rows reserved for post-training evaluation.

Forty real-world drug and medical-device safety claims: recalls, FDA Adverse Event Reporting System/MAUDE report counts, and current FDA-mandated boxed-warning/label text: each independently checked against openFDA, the FDA's own live structured-data API, across four distinct endpoints. A separate 1,000-row openFDA label proof evaluation has now been run through the browser-assisted ChatGPT workflow. Under strict deterministic grading, chatgpt-web answered 844/1,000 correctly; the 156 non-correct rows were concentrated in hard indications_and_usage label-section extraction. A 10,000-row remediation/training artifact and 1,000-row untouched held-out validation set now exist for that observed label-section weakness. No post-training improvement claim is made yet.

  • Every claim also tested live against Claude Sonnet 5 asked cold: 4 of 28 (14%) answers were confidently wrong on real drug-safety questions, exactly the kind of hallucination a health-AI product can't afford
  • Two structurally-identical-looking openFDA queries for the same drug/reaction pair returned different adverse-event counts (8,655 vs 8,652): a real, reproducible discrepancy between two common query shapes
  • Pulling a drug's full FDA label version history surfaced a genuine anomaly: a boxed warning removed by FDA action in 2016 briefly reappears in one intermediate label version, sandwiched between two versions that don't have it
  • A drug's current FDA label and independent adverse-event reporting corroborate the same safety signal from two completely different sources: the same reaction is the #1 most-reported FAERS term for that drug
40 rows · JSON · remediation
Metadata and sample rowsDiscuss this dataset

Drug Indication Diagnostic Archive (retired)

Drug-indication source-grounding data: graded diagnostic set, proof eval, 10k remediation data, and held-out eval

Retired from sale
CapabilityDrug indication source grounding
Designationmixed
Version2026-08-28-proof-backed-10k
Updated2026-08-28
SourcesNLM RxNav, RxClass, MED-RT DISEASE may_treat
FormatsJSON, JSONL eval export, JSONL remediation records
Product typeproduct family
Pricing modelcomponent pricing
338Diagnostic/eval rowsPublished product, fully model-graded
33.7%Measured hallucination rate114/338 cold answers hallucinated on the published diagnostic/eval set
500Fresh proof evalIndependent RxCUIs/questions, fully graded via browser-assisted ChatGPT run
78.4%Proof error rate392/500 incorrect or hallucinated; Wilson 95% CI 74.6%-81.8%
10,000Remediation data9,000 train / 500 dev / 500 test, validation passed
1,000Held-out evalUntouched RxNav rows reserved from remediation generation
Historical diagnostic export — retired from salediagnostic eval
Retired from sale

338 actual rows

Existing published JSON product with model-graded diagnostic/eval rows.

Request package
Drug Indication Proof Evalproof eval
$999 once

500 actual rows

Packaged 500 unique cases, provenance, historical grades, manifest and internal-use license. Automatic delivery after confirmed payment.

View checkout — $999
Drug Indication Remediation 10kremediation training
$4,900 reference price

10,000 actual rows

Validated training/remediation records with train/dev/test split and source-record leakage checks.

Request package
Held-Out Validationheld out validation
$1,250 reference price

1,000 actual rows

Untouched held-out RxNav evaluation set reserved from remediation generation.

Request package
Full Drug-Indication Improvement Packfull improvement pack
$6,500 reference price

11,500 actual rows

Proof eval, 10k remediation data, and held-out validation packaged together. No post-training improvement result is included yet.

Request package
Custom Model Failure Studycustom dataset
Custom quote

Customer-specific diagnostic, proof, remediation, and held-out study scoped to the buyer model/domain.

Request package
Published graded diagnostic/eval setdiagnostic
338 rows

historical archive; retired from sale

Fully model-graded; source-grounded against RxNav/RxClass MED-RT.

Fresh proof-evaluation pilotproof eval
500 rows

packaged for purchase with automatic delivery

Browser-assisted ChatGPT run: 108 correct, 392 incorrect/hallucinated, 0 refused, 0 ungraded.

Remediation/training datasetremediation
10,000 rows

validated locally; pending publication

Deterministic 9,000/500/500 train/dev/test split; factual answers derived from verified source rows.

Untouched held-out evaluationheld out
1,000 rows

validated locally; reserved from remediation

Zero RxCUI and row-ID overlap with the 10k remediation dataset.

Post-training improvement reportimprovement report
Pending

not yet run

Infrastructure exists, but no drug-indication post-training result has been measured.

A source-grounded drug-indication product family built from the National Library of Medicine RxNav/RxClass MED-RT DISEASE may_treat API. The published 338-row diagnostic/evaluation set is fully model-graded and found a measured 33.7% hallucination rate (114/338) when asked cold with no source access. A fresh 500-row proof evaluation was then run through the browser-assisted ChatGPT workflow and deterministically graded against RxNav/MED-RT, finding 392/500 incorrect or hallucinated answers. A separate 10,000-record remediation/training dataset and a 1,000-row untouched held-out evaluation now exist with source-record leakage checks. No post-training improvement claim is made yet; the held-out measurement step still needs an actual trained model run.

  • Measured cold-failure rate: 33.7% (114/338) across the published 338-row diagnostic/eval set: every row model-graded, not a partial sample
  • Fresh 500-row proof evaluation generated from live RxNav/RxClass MED-RT data and graded locally from browser-assisted ChatGPT JSONL responses: 108 correct, 392 incorrect/hallucinated
  • Separate 10,000-record remediation dataset with provenance, validation status, generation metadata, and train/dev/test split; not reused as final proof of improvement
  • The observed proof failure taxonomy is grounded in actual incorrect responses: over-specific clinical answers, unsupported indications, and wrong use/symptom scope
  • Ships with 392 DPO preference pairs derived from measured failures in the fresh proof run; use them as remediation data, never as held-out proof
338 rows · JSON · remediation
Metadata and sample rowsDiscuss this dataset
Verified evaluation catalog

Smaller verified datasets for evaluation and regression testing. These are not presented as full remediation systems unless proof/remediation assets exist.

Mathematical Claims Verification

Theorem, constant, and attribution claims checked by real computation, not recall

$299 reference price
FormatsJSON
Product typesmall verified eval
Pricing modelfixed one time

Small verified evaluation export; no separate proof/remediation asset is packaged for this product yet.

Forty-eight real mathematical claims (famous theorem statements, numeric sequence/constant properties, computational results, open-problem status checks, and "who first proved X" attributions), each checked against code actually written and run (Python sympy/mpmath, a from-scratch BigInt library, and an exhaustive brute-force combinatorial search), the real Wikidata knowledge graph, real published papers via OpenAlex, and, for one live current-events check, primary web sources fetched live. An AI model asked to check its own math claim just restates it more confidently; this dataset actually computes, queries, and looks the answer up, and says so honestly on the rare row where a second independent check genuinely doesn't exist. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Every claim also tested live against Claude Sonnet 5 asked cold, with no source access: 3 of 48 (6%) answers were confidently wrong, including misattributing 2^31− 1's primality proof to Édouard Lucas instead of the verified Euler
  • A second batch of 5 claims (a large Mersenne prime, a perfect number, a 200-digit Fibonacci value, a power-tower last-digit derivation, and a Catalan number) went 5-for-5 correct with full derivations shown: a genuine, unforced result, not cherry-picked
  • Re-audited every original row live and found one real self-correction: Wikidata does carry a discoverer statement for 2^31−1 being prime (Euler, 1772), on the specific-number entity, not the generic Mersenne-prime class entity checked the first time
  • Added an exhaustive brute-force proof that the Ramsey number R(3,3)=6 (all 32,768 edge-colorings of K6 checked by real code, every one contains a monochromatic triangle), and confirmed Wikidata's "credits the poser, not the solver" pattern on 3 more rows
  • Live-fetched today's actual largest known prime (2^136,279,841−1, GIMPS, Oct 2024) from two independent primary sources and found Wikidata's own record for the same fact frozen at a 1951–52 value: a real, current structured-data staleness gap

Scientific Claims Verification

Physical constants and discovery claims checked against NIST, OpenAlex, Wikidata, and real independent computation

$299 reference price
FormatsJSON
Product typesmall verified eval
Pricing modelfixed one time

Small verified scientific-factuality evaluation export.

Forty-three real claims about physical constants, first-discovery/first-measurement attributions, and specific published experimental results across physics, chemistry, astronomy, biology, and earth science, each checked against NIST's live CODATA fundamental-constants database, the actual paper reporting a result via OpenAlex, Wikidata's structured discoverer/date/value statements, or, new this pass, real independently executed computation deriving one fundamental constant from others. Fills the gap between Groundtruth Data's clinical-trial and citation-graph products, which don't touch general physical-science claims, and is now the catalog's most methodologically diverse verification product. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Every claim also tested live against Claude Sonnet 5 asked cold: 5 of 43 (12%) answers were confidently wrong, including fabricating a different 2015 Avogadro-constant figure than the real IAC-published value
  • Ran two from-scratch computational cross-checks this session, independently re-deriving the Rydberg constant and the c=1/√(μ₀ε₀) identity from five raw CODATA inputs each via real executed Python, agreeing with NIST's own published values to 12-13 significant figures
  • Found Wikidata's own unreferenced discovery-year statement for buckminsterfullerene (1984) contradicts the actual discovery paper's real publication date (November 1985), and found four disagreeing referenced river-length sources on Wikidata's own Amazon entity spanning nearly 600 km
  • Re-verified all 24 original rows live with zero drift, then added a genuinely independent second method to 7 rows, while explicitly declining to force a weak second-method addition where the only matching source didn't actually say what it needed to say

Patent & IP Claims Verification

Patent and IP claims checked against the live USPTO record, not what a company says about them

$249 reference price
FormatsJSON
Product typesmall verified eval
Pricing modelfixed one time

Small verified IP/factuality evaluation export.

Forty real claims companies and named inventors have made about patents ("we hold a patent on X," "patent number Y covers Z," "first to patent W"), each checked against the actual live record at USPTO's Patent Public Search system: does the patent exist, is it assigned to the claimed entity, do its real granted claims actually cover what's being claimed, and is it still valid today. Every row is now also cross-checked against a second, independent USPTO data feed: the live Patent Assignment and PTAB docket history tracking what actually happened to a patent after it was granted, not just its bibliographic snapshot: which caught patents that quietly lapsed for non-payment years after being cited in litigation, ownership chains that don't match a company's own marketing page, and three real errors in this dataset's own earlier release, corrected in place and disclosed rather than silently fixed. The original 23 rows also carry a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Expanded from 23 to 40 rows with a second independent verification method added to every row: USPTO's own live Patent Assignment and PTAB docket feed, not just a bibliographic lookup: which caught three real errors in the original release's own explanations and disclosed all three in place rather than silently rewriting them
  • Found a company's live product page still claiming present-tense patent protection under a specific patent number that USPTO's own record shows actually lapsed for non-payment more than five years ago
  • Traced Theranos's real patent portfolio surviving Elizabeth Holmes's fraud collapse: live USPTO assignment records independently confirm it was sold to a Fortress Investment Group shell company, which later used the same patents to sue an unrelated COVID-19 test maker

Clinical Trial Outcomes Verification

Does the trial data behind a drug claim actually back it up: checked against the real registry

from $349 reference price
Capabilityclinical trial registry grounding
Version2026-08-30
Updated2026-08-30
SourcesClinicalTrials.gov API v2, openFDA, Drugs@FDA, Purple Book, PubMed/OpenAlex where applicable
FormatsJSON, JSONL proof eval, JSONL remediation records, SFT JSONL, DPO JSONL, held-out eval JSONL
Product typeproduct family
Pricing modelcomponent pricing
47Legacy diagnostic rowsExisting published clinical-source verification export.
500Source-aware proof rowsFresh ClinicalTrials.gov API v2 proof rows graded against source truth.
9.0%Proof accuracy45 correct, 4 substantive incorrect, and 451 unnecessary refusal or uncertainty responses.
10,000Remediation rowsTargeted refusal-calibration and ClinicalTrials.gov fact-grounding rows.
1,000Held-out rowsUntouched held-out validation rows isolated from proof and remediation NCT IDs.
455Observed-failure DPO pairsRejected responses are actual proof-run refusals or wrong answers.

Legacy diagnostic export is fixed-price. Expanded proof, remediation, held-out, and DPO package should be quoted as a clinical refusal-calibration system.

Legacy clinical-trial exportdiagnostic
47 rows

published

Existing clinical-source verification product.

ClinicalTrials.gov proof evalproof eval
500 rows

source_aware_graded

Source-aware result: 45 correct, 4 substantive incorrect, and 451 unnecessary refusal or uncertainty responses.

ClinicalTrials.gov remediationremediation
10,000 rows

validated

Built to teach supported answers when ClinicalTrials.gov contains the fact and calibrated abstention when the source is insufficient.

ClinicalTrials.gov held-out validationheld out
1,000 rows

validated

Untouched held-out validation with zero proof or remediation NCT overlap.

Forty-seven real claims about clinical trial outcomes and FDA drug approvals remain available as the legacy diagnostic export. The expanded ClinicalTrials.gov Fact Grounding and Refusal Calibration package adds a 500-row proof evaluation, a 10,000-row remediation set, a 1,000-row untouched held-out set, and 455 observed-failure DPO pairs. The source-aware proof result showed 9.0% accuracy with 451 unnecessary refusals or uncertainty responses on directly verifiable ClinicalTrials.gov facts and 4 substantive wrong answers. This is framed as refusal calibration and registry fact grounding, not as a fabricated-fact hallucination benchmark.

  • Every claim also tested live against Claude Sonnet 5 asked cold: 9 of 47 (19%) answers were confidently wrong, including two cases where the model correctly recalled a real drug approval but wrongly guessed it was filed as a supplement rather than an entirely separate FDA application
  • Two different drugs: bardoxolone methyl and sotorasib (Lumakras): met their confirmatory trial's own statistically significant primary endpoint and were still rejected or held back by FDA, confirmed live against both the trial registry's structured statistics and openFDA's actual submission history
  • 43 of 47 rows now carry two or more independent live-checked sources per claim (trial registry, FDA approval record, a second FDA database, and/or the actual publication): including a real transcription error this pass found in a prior release's Aduhelm/aducanumab row and fixed with the correction disclosed in place
  • The source-aware 500-row ClinicalTrials.gov proof run found 451 unnecessary refusals or uncertainty responses on verified registry facts, plus 4 substantive wrong answers
  • The validated remediation package contains 10,000 rows across trial status, phase, enrollment, sponsor, completion date, results availability, locations, conditions, primary outcomes, interventions, and eligibility

Government Contract Verification

Federal contracting claims checked against the actual USASpending.gov award record

$249 reference price
FormatsJSON
Product typesmall verified eval
Pricing modelfixed one time

Small verified government-contract evaluation export.

Forty real claims companies make about their US federal and defense contracting relationships: "prime DoD contractor," specific dollar figures, named agency relationships: each checked against the real award data on USASpending.gov, the federal government's own system of record for who it has actually paid, with a second independent method (real SEC EDGAR filings) added wherever one genuinely exists. Surfaces real, checkable failure modes: contract-ceiling figures presented as if already earned, a real NYSE-listed company with billions in current federal work that's entirely invisible under its own name because every dollar is still filed under its pre-merger legal entity, and: new this pass: a household-name defense prime whose federal business is invisible to a simple name search because the identical legal entity is registered under more than 100 separate, un-consolidated USASpending IDs. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Added a second, independent verification method (real SEC EDGAR filings, not just USASpending) to 8 rows: e.g. quantifying that CrowdStrike's audited $4.8B SEC-reported revenue is over 24,000x its entire confirmed direct-prime federal award history, and independently confirming via a wholly different accounting concept that 83.6% of Lockheed Martin's audited net sales trace to federal contracts
  • A live re-audit of the original 25 rows found zero errors but surfaced a brand-new failure mode: Northrop Grumman's own legal name is registered as 158 separate, un-consolidated USASpending recipient IDs, undercounting its real $42B-revenue federal business by nearly 70% under a naive name search
  • Every claim also tested live against Claude Sonnet 5 asked cold: only 5 of 40 answered correctly with real confidence, 3 were confidently wrong (hallucinated), and 31 honestly hedged rather than guessing at specific post-cutoff contract figures

Web-Navigation Trajectories

Real browser sessions with captured structural data, including real failures

$249 reference price
FormatsJSON
Product typesmall verified eval
Pricing modelfixed one time

Small verified browser-trajectory export.

Forty-two real browser-use sessions: an AI system actually navigating, clicking, typing, hovering, and scrolling a live browser against real public websites, with the exact action sequence, the real accessibility-tree/selector data captured at each step, and the real fact or outcome found at the end: plus, on a dozen rows, a second genuinely independent verification method (Wikidata's structured API, live REST API calls, raw GitHub source, NIH PubChem, caniuse.com) layered on top of the original browser session. Includes multi-hop tasks with genuine runtime branching, multi-tab flows, hover-triggered menus, infinite scroll, a fully unplanned Special:Random navigation whose destination was unknown before the session ran, and five honestly-recorded failures and partial-failures (CAPTCHA walls, login walls, a bot-mitigation block, a partial content gate), never dropped to make the success rate look better.

  • A live re-audit caught a real fact change in the wild: GitHub's #1 trending COBOL repository swapped entirely (a 3,588-star repo to a 253-star one) between two live captures taken a day apart, disclosed in place rather than silently overwritten
  • 12 originally single-source rows deepened with a genuinely independent second method (Wikidata, live REST API headers, raw CPython/PEP GitHub source, NIH PubChem): one traced a Wikidata citation to a real GitHub commit and found Rust's exact creation date, July 23, 2006, which appears nowhere in Wikipedia's own prose
  • New rows include a genuinely unplanned Wikipedia Special:Random navigation and a third real access pattern beyond clean success/failure: Instagram renders a real public follower count for an unauthenticated visitor but gates the actual content behind a forced sign-up modal

Fact-Verification Dataset

Real claims, two independent sources each, honest verdicts including disagreement

$499 reference price
FormatsJSON
Product typebroad verified eval
Pricing modelfixed one time

Broad verified factuality evaluation export with 237 rows.

Two hundred thirty-seven real-world claims across AI/tech, business, science, medicine, history, geography, culture, sports, entertainment, law, and government, each checked against a real cited source AND a second, genuinely independent source, with the two sources' agreement recorded explicitly. Built specifically to probe plausible-sounding falsehoods and implausible-sounding truths, not just easy trivia: and to surface the cases where two independent sources actually disagree, which a single-source lookup would never catch. A 2026-07-12 expansion added 37 new claims leaning into current events genuinely beyond any model's training cutoff (Nvidia's fiscal-2026 earnings, the 2026 Winter Olympics medal table, Super Bowl LX, the 98th Academy Awards), verified live via real API calls, Wikidata SPARQL queries, and independent web sources, alongside a full re-audit of the original 200 that found the existing data held up with zero new errors. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Expanded from 200 to 237 claims with 37 new dual-sourced rows on 2026 events verified live rather than recalled: Nvidia's $215.9B fiscal-2026 revenue, the 2026 Winter Olympics medal table, Super Bowl LX, and the 98th Academy Awards
  • A live re-audit of all 8 flagged source disagreements plus every AI/tech claim found the original 200 rows held up with zero new errors, while a live SEC EDGAR query caught Nvidia silently switching its revenue XBRL tag after fiscal 2022: the same 'coverage gap, not staleness' pattern already documented for Wikidata
  • Model tested cold across all 237 claims: 214 correct, 16 confidently wrong (6.8%), and 7 honest hedges concentrated almost entirely in the new 2026 rows, where the model correctly declined to guess outcomes beyond its knowledge cutoff instead of fabricating an answer
237 rows · JSON · broad eval
Metadata and sample rowsDiscuss this dataset

Code-Reality Verification Dataset

Does this library API claim actually match what shipped: checked against the real published artifact

$299 reference price
FormatsJSON
Product typesmall verified eval
Pricing modelfixed one time

Small verified code/package factuality evaluation export.

Forty-two real claims about specific npm, PyPI, and (new this release) crates.io package APIs at exact versions, each checked by actually downloading and inspecting the real published tarball/sdist/crate file: and, for 38 of the 42 rows, confirmed a second, independent way by actually installing the real package and running the real code (or, for two Rust API-removal claims, capturing a real compiler error from trying to build against the claimed symbol). Targets the exact failure mode of hallucinated library APIs: functions, parameters, and behaviors that sound plausible but were never shipped, shipped differently, or existed only in a different version than the one claimed. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Every claim also tested live against Claude Sonnet 5 asked cold: 5 of 42 (12%) answers were confidently wrong on real npm/PyPI/crates.io package API questions
  • 38 of 42 rows independently confirmed two separate ways: real source inspection plus actually running the installed code, or for two Rust rows, a real compiler error: not just a text match
  • New crates.io ecosystem plus real dynamic execution caught a live nuance: on current Node.js, require()-ing an ESM-only package like chalk or node-fetch no longer throws outright: it silently returns a broken module object instead, catchable only by actually running it

UK Companies House Verification Dataset

Director, incorporation, status, and beneficial-ownership claims checked against the real UK company register

$349 reference price
FormatsJSON
Product typesmall verified eval
Pricing modelfixed one time

Small verified corporate-registry evaluation export with authenticated-source verification.

55 real claims about UK-registered companies (mostly AI/tech firms): director identity, incorporation dates, company status, registered addresses, beneficial ownership (PSC), cross-company appointment history, previous company names, SIC industry classification, filing history, and registered charges: each independently checked against the live UK Companies House register via both its free public website and its authenticated REST API. 82% of rows are backed by two genuinely independent, separately-coded verification paths; the rest are honestly single-sourced because no second method exists (PSC and appointment data have no public-HTML equivalent at all). Includes a real bug found and fixed in the dataset's own verification tooling this release, plus live model-graded testing on every row: asked cold with no tools, Claude Sonnet 5 hedges correctly on 95% of these claims and hallucinates a specific wrong answer on 2%, showing why a live register lookup beats an LLM's memory for this kind of fact.

  • Famous founders are repeatedly absent from their own UK subsidiary's board: Dario Amodei isn't a director of Anthropic Limited, Elon Musk isn't a director of XAI UK Limited, and Michael Truell (Anysphere/Cursor) isn't a director of Anysphere UK Ltd: the same real pattern confirmed a third time this release.
  • Coincidental company-name collisions are a systematic, repeatable trap, not a one-off: XAI UK Limited (2012) and Groq UK Limited (2004) both predate the AI companies they sound like by a decade or more, and 'Signal AI Ltd' (incorporated Nov 2025) is a completely unrelated company: the real Signal AI is legally registered as 'Signal Media Limited.'
  • This release found and fixed a real bug in its own verification code: a name-matching helper required exact-token equality, so ordinary claims like 'Vishal Marria' silently failed to match Companies House's fuller registered name 'MARRIA, Vishal Kumar': caught, root-caused, fixed, and re-verified live, not papered over.

AI Safety / Refusal Testing Dataset

Cross-model-tier refusal divergence and consistency, not a single one-shot test

$249 reference price
FormatsJSON
Product typepilot verified dataset
Pricing modelfixed one time

Pilot-scale verified refusal-testing export.

Thirty real prompts drawn from two published, MIT-licensed AI-safety benchmarks (JailbreakBench, HarmBench), each individually screened against a strict category allowlist (several on-allowlist-labeled candidates, including a nuclear-weapon prompt mislabeled "Government decision-making," were excluded on individual review) and tested fresh against three model capability tiers with repeated sampling per tier (180 total live invocations), not a single one-shot test against one model. Every row now also carries a second, independent verification method: a live re-fetch of the source benchmark's own raw data confirming an exact prompt/category match: after a full re-audit of the original 21 rows found zero drift. Strictly scoped to lower-severity categories only: nothing touching weapons, hacking, violence, sexual content, trafficking, or self-harm: and stores only classifications, never generated harmful content.

  • Every row now carries two independent verification methods: a live 3-tier x 2-sample behavioral test (180 cells) plus a live source-fidelity re-fetch confirming the prompt/category exactly matches the original benchmark's own data file
  • Individual per-prompt screening actively excludes on-allowlist-labeled prompts that fail deeper review: this pass caught and dropped a nuclear-weapon-construction prompt filed under "Government decision-making," an organ-trafficking prompt, and three others, rather than trusting the source benchmark's own category labels
  • Still only one tier-divergent prompt in 30: an investment-advice bot request where haiku refuses, sonnet reframes into a disclaimer-heavy tool, and opus (with agentic tool access) locates and ships a working artifact: capability going up, caution going down, for this one category only

Cross-Source Verification Dataset

Two independent real sources, checked against each other, not just one

$299 reference price
FormatsJSON
Product typesmall verified eval
Pricing modelfixed one time

Small verified cross-source consistency evaluation export.

Wikidata vs. live web; SEC filings vs. marketing claims; and, new in this expansion, one acquirer's audited SEC filing vs. another company's widely-reported deal-announcement figure for the same transaction. 42 real cases checking whether a structured official source actually agrees with a second, independent real source: and, where possible, a third. This pass re-verified all 27 original rows live with zero errors found, added a genuine second SEC-filed source to two rows that originally hit Wikidata's total silence (surfacing real gaps between announced and audited deal sizes), and expanded into 15 new rows covering newer AI companies and real corporate M&A, including a materially wrong headline acquisition-price figure caught by comparing an acquirer's own audited books against press coverage.

  • 42 rows across 3 methods (Wikidata-vs-live, marketing-vs-SEC-filing, and a new acquirer-filing-vs-deal-headline cross-check); the 2026-07-14 adversarial audit re-derived all 42 rows live, found every structured SEC/Wikidata value exact, and fixed 10 claimed-side citation errors in place with disclosed correction notes
  • New acquirer's-own-10-K vs press-release deal-price checks caught a real, material $1.03B (16%) gap for IBM's HashiCorp acquisition ($7.433B audited total consideration vs. the widely-repeated $6.4B headline) and resolved two prior Wikidata coverage gaps (Meta/Scale AI, CoreWeave/Weights & Biases) by finding the SEC-filed figures diverge from press-reported deal sizes
  • 7 new company rows (Perplexity, Groq, Character.AI, Cursor/Anysphere, Harvey AI, etc.) found Wikidata's ownership/CEO coverage gaps concentrate hardest on newer, fast-moving AI startups: 5 of 7 have zero structured statement at all, not just a stale one, including a live, unclosed $60B SpaceX-Cursor acquisition Wikidata hasn't recorded either way

Historical Event/Date/Figure Facts Verification

Historical event, date, and attribution claims independently checked against Wikidata's live structured statements, not model memory

$249 reference price
FormatsJSON
Product typesmall verified eval
Pricing modelfixed one time

Small verified history factuality evaluation export.

Forty-two real historical claims: a battle date, an assassination's causal framing, who really invented the telephone: each checked against Wikidata's real, live, CC0-licensed structured statements for the specific entity in question, including a purpose-built technique for compound multi-date events (space missions, voyages) where a naive property lookup would silently grab the wrong sub-event's date. most claims match, a documented minority diverge (off-by-one-day/year errors, invention misattributions, an overstated single-cause claim), and 1 is honestly marked unverifiable. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Every claim also tested live against Claude Sonnet 5 asked cold: 0 of 34 answers were wrong outright, the strongest clean-recall performance found anywhere in this catalog
  • Invention-attribution folklore fails twice, independently: Wikidata lists Reis, Gray, Bell, and Meucci for the telephone (no Edison at all), and five credited inventors for the incandescent light bulb
  • Three genuine "almost right" date errors caught by checking the exact claimed day, not just the year: the Titanic struck the iceberg April 14 but sank April 15; Lincoln was shot April 14 but died April 15
  • Wikidata's own "has cause" statement for World War I resolves to an item literally labeled "multiple causes," directly contradicting the popular single-trigger claim that Franz Ferdinand's assassination alone caused the war

English Language & Literary Attribution Verification

Famous quotes and book-publication claims checked against real public-domain full text and live Wikidata, not recalled

$249 reference price
FormatsJSON
Product typesmall verified eval
Pricing modelfixed one time

Small verified literary attribution evaluation export.

Forty-two claims, each traced back to a real, live-fetched source: exact substring search against the real full plain-text of 25 public-domain books/plays via Project Gutenberg for quote claims, and live Wikidata publication-date/authorship checks. 16 of the original 35 rows (46%) are genuine divergences: misquotes, mixed-up authors, and swapped years that an LLM's training-data recall would very plausibly get wrong with high confidence: a count that grew by one when this dataset's own adversarial audit held a row to its own punctuation-strict standard and flipped it. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Every claim also tested live against Claude Sonnet 5 asked cold: 5 of 35 (14%) answers were confidently wrong, on top of the 16 divergences already found in the underlying source claims
  • "Elementary, my dear Watson" does not appear anywhere in Arthur Conan Doyle's original canon: confirmed by full-text-searching all 9 real Doyle books on Gutenberg, not just one
  • Oscar Wilde's "I can resist everything except temptation" is real but misattributed: not in The Picture of Dorian Gray at all, but verbatim in his play Lady Windermere's Fan
  • A live Wikidata check catches two planted plausible-sounding errors cleanly: The Great Gatsby claimed as the same year as Ulysses (1922; real value 1925), and Wuthering Heights claimed as Charlotte Brontë's (real author: her sister Emily)

Geographic Facts Verification

Country capitals, populations, land areas, borders, and physical-geography superlatives checked against two independent live sources

$249 reference price
FormatsJSON
Product typesmall verified eval
Pricing modelfixed one time

Small verified geography factuality evaluation export.

Forty real geographic claims: capitals, population figures with year, land areas, bordering countries, and physical-geography superlatives: each cross-checked against two live, independently maintained sources: the CIA World Factbook's last public snapshot (public domain, US government work) and live Wikidata. Surfaces genuinely current, non-obvious findings, including a stale Wikidata population statement that would get "which country is more populous" wrong, and a currently unsettled naming dispute over North America's highest peak. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Every claim also tested live against Claude Sonnet 5 asked cold: 2 of 30 (7%) answers were confidently wrong
  • Wikidata's population statement for India is stuck at a 2020 figure while China's is current to 2025: a tool trusting Wikidata alone would wrongly show China still ahead, when fresher 2025 estimates show India has actually overtaken China
  • Wikidata's entity for North America's highest point currently has no English label at all, reflecting a live, unsettled dispute after a 2025 US federal order reverted the peak's official name
  • The CIA World Factbook's official site was itself taken offline by the CIA in Feb 2026: confirmed live via redirect to a "farewell" notice: forcing this product to source from the last public snapshot via a long-running open mirror

Language & Runtime Semantics Verification

Claims about Python/JavaScript runtime behavior, each settled by actually running the code or fetching the official changelog

$299 reference price
FormatsJSON
Product typesmall verified eval
Pricing modelfixed one time

Small verified runtime-semantics evaluation export with an improvement-report scaffold.

Forty-six claims about sort stability, integer/float precision, division/modulo sign semantics, string encoding edge cases, mutability/identity gotchas, and "which version introduced feature X," verified by literally executing python3/node and capturing the real output, or by live-fetching the language's own official changelog. Most rows carry real dual-interpreter executed code with literal output; the rest carry a live-fetched changelog snippet. Includes several deliberately adversarial rows built from real, commonly-stated wrong claims, to prove the method actually catches errors. Every row also carries a live test: the same fact, asked as an open question cold to Claude Sonnet 5 with no source access, graded against the verified answer, turning this from an answer key into a graded exam.

  • Every claim also tested live against Claude Sonnet 5 asked cold: 2 of 46 (4%) answers were confidently wrong on live-executed runtime-behavior questions
  • A second batch of 6 claims (string immutability, list identity vs. equality, dict insertion order, JS type-coercion quirks, shared-reference list multiplication, and a Python 3.12 syntax changelog fact) went 6-for-6 correct: a genuine, unforced result
  • Node.js's global fetch() shipped in 18.0.0 explicitly marked "experimental" and wasn't promoted to stable until 21.0.0: a real 3-major-version gap a compressed claim like "Node 18 shipped stable fetch" gets checkably wrong
  • Running 5 % -3 in both live interpreters shows Python's % takes the sign of the divisor while JavaScript's % takes the sign of the dividend: same operands, silently different answers
  • Python's `nan in [nan]` returns True even though `nan == nan` is False, because CPython's list-membership test tries object identity before falling back to equality: isolated by testing against a second, genuinely different nan object

Tier 3: Custom Model Failure Study

Start with a suspected weakness and target capability; Groundtruth Data scopes diagnostic, proof, remediation, and held-out assets around it.

Request custom study

Looking for the company landscape archive?

171 companies across 14 categories, sourced and cited.

Ask about this archive