Evaluation pilots
Find the factual failures that matter to your team.
Measure a defined capability against source-backed reference answers. Inspect the model output, evidence, and grading rule behind each result.
How an evaluation works
Scope the task
Agree on questions, source coverage, and what counts as a correct answer.
Build the proof set
Preserve reference answers and provenance, including unsupported cases.
Run and grade
Retain model outputs and apply the agreed rules.
Review the failures
Inspect patterns and choose a focused next step.
What you receive
A baseline, evidence-backed failure examples, and the agreed export or report. Remediation and follow-up validation are separately scoped.
Custom evaluations are operator-led. We agree on price and turnaround before work starts.
A recorded example
Our SEC proof run contains 500 cases: 136 correct and 364 incorrect, a 27.2% baseline. This is our own task-specific evaluation, not an independent or universal model score.
Read the study and its limits →