Built to help models answer source-supported PubMed metadata while preserving appropriate abstention when metadata is genuinely unavailable.
The proof evaluation asked PubMed citation metadata questions covering PMID identity, DOI recall, paper title, journal, publication date, first author, author lists, paper metadata, paper existence, and retraction or correction classification.
The primary source of truth is NCBI PubMed metadata. Rows preserve PMID, DOI where available, title, journal, publication date, authors, publication types, source URLs, and provenance.
On a 500-row PubMed-backed proof evaluation, the tested model achieved 32.2% accuracy. 305 of 339 non-correct responses were unnecessary refusal or uncertainty on answerable PubMed metadata. This is presented as PubMed metadata grounding and refusal calibration.
305 rows
33 rows
1 row
29 of 40 rows
The strongest refusal patterns were PMID identity, DOI recall, paper title, journal, publication date, first author, and author-list questions. The proof also identified 29 substantive errors across 40 retraction/correction classification questions.
Groundtruth Data built 9,573 validated remediation rows. The rows target answerable PubMed metadata, careful DOI and author identity, retraction/correction classification, and calibrated abstention when the source does not provide enough information.
No claim is made that this remediation has improved a model. Improvement must be measured in a separate before and after experiment using untouched held-out data.
The DPO file contains 339 real observed proof-failure pairs. It includes 305 refusal or uncertainty-derived pairs and 34 substantive wrong-answer-derived pairs. No rejected answers were fabricated.
A separate 970-row PubMed held-out set is reserved for post-training evaluation. It is not used for proof scoring, remediation generation, prompt tuning, training, or checkpoint selection.
Teams use the proof to measure a baseline, the remediation split for training or fine-tuning, the DPO pairs for preference experiments when appropriate, and the held-out set for final measurement. The held-out rows remain separate from training.
500 actual rows
500-row PubMed-backed proof evaluation with source-aware deterministic grading.
Request package9,573 actual rows
Validated remediation rows targeting observed refusal, metadata identity, and retraction/correction failures.
Request package970 actual rows
Untouched held-out rows excluded from proof and remediation PMIDs, DOIs, source records, and source URLs.
Request package11,043 actual rows
500-row proof eval, 9,573 remediation rows, and 970 held-out validation rows. No model improvement result is included.
Request packageAny larger PubMed package requires separate source-supply and validation review.
Request packageWe can package the proof evaluation, 9,573 remediation rows, observed-failure DPO pairs, and held-out validation set for your training or evaluation workflow.