Evaluation cases
Questions, reference answers, and grading rules tied to a defined task. Keep answers out of the prompts you send to the model.
Groundtruth Data
Use source-backed cases to find factual failures, inspect why an answer is wrong, and define the data needed to address it.
How it works
A proof evaluation measures the starting point. Remediation and validation are separate stages, not promises of automatic improvement.
Choose the question your system needs to answer and the sources that can establish correctness.
Run a source-backed evaluation, retain model outputs, and inspect the evidence behind each grade.
Use verified corrections to scope remediation data. Your team controls changes to its model or system.
Test on separate cases. Report what changed, including regressions and unresolved errors.
What you receive
Questions, reference answers, and grading rules tied to a defined task. Keep answers out of the prompts you send to the model.
Source references and retained evidence. Recorded model outputs and grades are included where the package specifies them.
Structured files, a manifest, and explicit usage terms. Review the schema, source dates, and limitations before use.
Packaged dataset
A historical MED-RT source-grounding evaluation: 500 unique questions, verified reference answers, source provenance, and recorded model grades.
This is separate from the retired 338-row diagnostic export. It is not clinical advice, an independently blinded holdout, or evidence of model improvement.
Browse the catalog
Explore the existing inventory by task. The 500-case drug proof package above has automatic delivery. Other datasets require a reviewed scope and delivery agreement; larger remediation and validation assets remain pending publication.
Reference counts describe the listed export, not the combined size of every component in a product family. Component counts, source dates and limitations are in the dataset details.
Full catalog and component pricing →Custom projects
A scoped pilot can combine a baseline evaluation, evidence-backed failure examples, and targeted remediation data. Separate validation measures the result of an intervention your team makes.
Discuss a data pilotLarger remediation and held-out inventories require a reviewed package. They are not included in the 500-case checkout or automatically available for download.
Customer prompts, outputs, and failure records stay private by default. We do not silently train or modify your model.
Why the datasets stay separate →