Groundtruth Data

Evaluation data with evidence behind every answer.

Use source-backed cases to find factual failures, inspect why an answer is wrong, and define the data needed to address it.

How it works

Four steps, one defined capability.

A proof evaluation measures the starting point. Remediation and validation are separate stages, not promises of automatic improvement.

  1. Define the task

    Choose the question your system needs to answer and the sources that can establish correctness.

  2. Measure the baseline

    Run a source-backed evaluation, retain model outputs, and inspect the evidence behind each grade.

  3. Target the failures

    Use verified corrections to scope remediation data. Your team controls changes to its model or system.

  4. Evaluate the change

    Test on separate cases. Report what changed, including regressions and unresolved errors.

What you receive

Data your team can inspect and use.

Evaluation cases

Questions, reference answers, and grading rules tied to a defined task. Keep answers out of the prompts you send to the model.

Provenance and results

Source references and retained evidence. Recorded model outputs and grades are included where the package specifies them.

A reproducible export

Structured files, a manifest, and explicit usage terms. Review the schema, source dates, and limitations before use.

Packaged dataset

500-case Drug Indication Proof Evaluation

A historical MED-RT source-grounding evaluation: 500 unique questions, verified reference answers, source provenance, and recorded model grades.

  • JSONL evaluation and provenance files
  • Historical model responses and grades
  • SHA-256 manifest and internal-use license

This is separate from the retired 338-row diagnostic export. It is not clinical advice, an independently blinded holdout, or evidence of model improvement.

Browse the catalog

21 dataset and product-family entries.

Explore the existing inventory by task. The 500-case drug proof package above has automatic delivery. Other datasets require a reviewed scope and delivery agreement; larger remediation and validation assets remain pending publication.

Dataset / product familyListed reference rowsPurchase pathExplore
Mathematical Claims Verification48Request scope and delivery quoteView sample
Scientific Claims Verification43Request scope and delivery quoteView sample
Patent & IP Claims Verification40Request scope and delivery quoteView sample
Legal Citation & Case Law Verification42Request scope and delivery quoteView sample
Clinical Trial Outcomes Verification47Request scope and delivery quoteView sample
Government Contract Verification40Request scope and delivery quoteView sample
PubMed Citation Metadata and Refusal Calibration9,573Request scope and delivery quoteView sample
Web-Navigation Trajectories42Request scope and delivery quoteView sample
Fact-Verification Dataset237Request scope and delivery quoteView sample
Code-Reality Verification Dataset42Request scope and delivery quoteView sample
Citation / OpenAlex Metadata Grounding43Request scope and delivery quoteView sample
UK Companies House Verification Dataset55Request scope and delivery quoteView sample
AI Safety / Refusal Testing Dataset30Request scope and delivery quoteView sample
Cross-Source Verification Dataset42Request scope and delivery quoteView sample
Historical Event/Date/Figure Facts Verification42Request scope and delivery quoteView sample
English Language & Literary Attribution Verification42Request scope and delivery quoteView sample
Geographic Facts Verification40Request scope and delivery quoteView sample
Language & Runtime Semantics Verification46Request scope and delivery quoteView sample
SEC / EDGAR Financial Grounding41Request scope and delivery quoteView sample
FDA Drug/Device Safety Verification40Request scope and delivery quoteView sample
Drug Indication Diagnostic Archive (retired)338Historical diagnostic retired; 500-case package aboveView sample

Reference counts describe the listed export, not the combined size of every component in a product family. Component counts, source dates and limitations are in the dataset details.

Full catalog and component pricing →

Custom projects

Build around a real failure in your system.

A scoped pilot can combine a baseline evaluation, evidence-backed failure examples, and targeted remediation data. Separate validation measures the result of an intervention your team makes.

Discuss a data pilot

Scope before scale

Larger remediation and held-out inventories require a reviewed package. They are not included in the 500-case checkout or automatically available for download.

Your data stays yours

Customer prompts, outputs, and failure records stay private by default. We do not silently train or modify your model.

Why the datasets stay separate →