← All articles

Evidence study

500 SEC filing questions: what a 27.2% baseline tells us

A recorded Groundtruth evaluation, its failure patterns, and the limits of the headline number.

Related reading: our original Medium post. This company-site edition reflects the current product scope.

The result and its scope

Groundtruth built this evaluation and reports its own findings. This is not an independent benchmark or customer testimonial. The recorded browser-assisted chatgpt-web run contains 500 SEC financial grounding cases: 136 correct, 364 incorrect, no refusals, and no ungraded cases. Baseline accuracy was 27.2%.

That result belongs to this task set and recorded run. It does not establish the performance of every model, identify a universally applicable error rate, or demonstrate a post-training improvement.

What the questions checked

The proof set ties financial facts and calculations to SEC filing and XBRL provenance. A reference answer depends on the company, concept, reporting period, units, and relevant filing. A plausible value from a different period can still be wrong.

The recorded failure analysis includes wrong numeric values, temporal confusion, wrong dates, and calculation errors. These families can overlap; their counts should not be added as though each incorrect answer belonged to exactly one category.

Why the reference needs an audit trail

The source reference and grading rule are as important as the model response. Exact-value, date, and tolerance-based checks answer different questions. Without preserving those distinctions, a benchmark error can look like a model error.

The public 75-row JSONL sample contains questions, verified answers, recorded model answers, grades, failure types, verification summaries, and source URLs. It offers a way to inspect individual cases rather than rely on the aggregate alone.

What remains to be demonstrated

The repository separates proof evaluation, remediation, and held-out assets. The larger assets remain subject to commercial packaging and review. Their existence does not show that a model has improved.

A defensible follow-up records the intervention, model and system settings, overlap checks, and results on separate evaluation cases. Until that experiment is complete, this study establishes a baseline and observed failure patterns. It does not establish a successful fix.