Back to products
Product family

PubMed Citation Metadata and Refusal Calibration

Built to help models answer source-supported PubMed metadata while preserving appropriate abstention when metadata is genuinely unavailable.

Request PubMed packageDownload public sample
What Was Tested

The proof evaluation asked PubMed citation metadata questions covering PMID identity, DOI recall, paper title, journal, publication date, first author, author lists, paper metadata, paper existence, and retraction or correction classification.

CapabilityPubMed citation metadata grounding and refusal calibration
Designationmixed
Version2026-08-31
Updated2026-09-01
SourcesNCBI PubMed E-utilities
FormatsJSONL records, JSONL eval, JSONL SFT, JSONL observed-failure DPO, held-out eval JSONL
Product typeproduct family
Pricing modelcomponent pricing
Source Of Truth

The primary source of truth is NCBI PubMed metadata. Rows preserve PMID, DOI where available, title, journal, publication date, authors, publication types, source URLs, and provenance.

Proof Result
500Proof rows
32.2%Accuracy
67.8%Non-correct
305Refusal or uncertainty
34Substantive incorrect

On a 500-row PubMed-backed proof evaluation, the tested model achieved 32.2% accuracy. 305 of 339 non-correct responses were unnecessary refusal or uncertainty on answerable PubMed metadata. This is presented as PubMed metadata grounding and refusal calibration.

Observed Failure Modes
unnecessary refusal or uncertainty

305 rows

wrong entity

33 rows

false citation

1 row

retraction and correction errors

29 of 40 rows

The strongest refusal patterns were PMID identity, DOI recall, paper title, journal, publication date, first author, and author-list questions. The proof also identified 29 substantive errors across 40 retraction/correction classification questions.

What Groundtruth Built

Groundtruth Data built 9,573 validated remediation rows. The rows target answerable PubMed metadata, careful DOI and author identity, retraction/correction classification, and calibrated abstention when the source does not provide enough information.

No claim is made that this remediation has improved a model. Improvement must be measured in a separate before and after experiment using untouched held-out data.

DPO

The DPO file contains 339 real observed proof-failure pairs. It includes 305 refusal or uncertainty-derived pairs and 34 substantive wrong-answer-derived pairs. No rejected answers were fabricated.

Held-Out Validation

A separate 970-row PubMed held-out set is reserved for post-training evaluation. It is not used for proof scoring, remediation generation, prompt tuning, training, or checkpoint selection.

Data Diversity
9,145Remediation PMIDs
9,122Remediation DOIs
2,467Remediation journals
970Held-out PMIDs
920Held-out DOIs
553Held-out journals
How Teams Use It

Teams use the proof to measure a baseline, the remediation split for training or fine-tuning, the DPO pairs for preference experiments when appropriate, and the held-out set for final measurement. The held-out rows remain separate from training.

Pricing
PubMed Metadata Proof Evalproof eval
$1,299 reference price

500 actual rows

500-row PubMed-backed proof evaluation with source-aware deterministic grading.

Request package
PubMed Metadata Remediationremediation training
$6,900 reference price

9,573 actual rows

Validated remediation rows targeting observed refusal, metadata identity, and retraction/correction failures.

Request package
PubMed Held-Out Validationheld out validation
$1,500 reference price

970 actual rows

Untouched held-out rows excluded from proof and remediation PMIDs, DOIs, source records, and source URLs.

Request package
PubMed Metadata Calibration Packfull improvement pack
$8,500 reference price

11,043 actual rows

500-row proof eval, 9,573 remediation rows, and 970 held-out validation rows. No model improvement result is included.

Request package
PubMed expansioncustom dataset
Custom quote

Any larger PubMed package requires separate source-supply and validation review.

Request package
Limitations

Request the PubMed package

We can package the proof evaluation, 9,573 remediation rows, observed-failure DPO pairs, and held-out validation set for your training or evaluation workflow.

Contact Groundtruth Data