Groundtruth Data found a repeatable weakness in period-specific SEC financial fact grounding. The strongest failures were numerical financial facts, fiscal dates, temporal comparisons, and YoY calculations.
The proof evaluation asked period-specific questions about SEC financial facts from companyfacts, companyconcept, submissions metadata, and 10-K and 10-Q filing metadata. Tasks covered revenue, net income, operating income, cash flow, cash and equivalents, assets, liabilities, acquisition values, fiscal year dates, period comparisons, and YoY calculations.
Every answer traces to SEC EDGAR structured data or a deterministic calculation from SEC values. Numeric rows preserve CIK, company, accession, form, filing date, fiscal period, XBRL concept, raw SEC value, unit, normalized verified answer, and source URL.
The observed error rate has a 95% Wilson confidence interval of 68.7% to 76.5%. This result is specific to the completed chatgpt-web proof run and deterministic SEC-source grading.
327 rows
65 rows
37 rows
30 rows
Identity tasks performed strongly and are not the focus of the remediation package.
Groundtruth Data built 10,000 validated remediation rows targeting the observed SEC failure modes. The remediation data is designed to help teams train against period-specific numeric grounding, date grounding, temporal comparisons, and deterministic financial calculations.
No claim is made that this remediation has improved a model. Improvement must be measured in a separate before and after experiment using untouched held-out data.
A separate 1,000-row SEC held-out set is reserved for post-training evaluation. It is not used for proof scoring, remediation generation, prompt tuning, training, or checkpoint selection.
41 actual rows
Existing published JSON export with verified SEC financial-filing rows.
Request package500 actual rows
Fresh SEC EDGAR proof set with completed chatgpt-web browser run and deterministic source-aware grading.
Request package10,000 actual rows
Validated SEC companyfacts/companyconcept remediation records targeting observed numeric, temporal, and calculation failures.
Request package1,000 actual rows
Untouched SEC EDGAR validation rows excluded from proof and remediation source records.
Request package11,500 actual rows
500-row proof eval, 10,000 remediation rows, and 1,000 held-out validation rows. No model improvement result is included yet.
Request packageRequires a separate source-supply and validation review before any larger package is claimed.
Request packageWe can package the proof evaluation, 10k remediation data, observed-failure DPO pairs, and held-out validation set for your training or evaluation workflow.