Evaluation pilots

Find the factual failures that matter to your team.

Measure a defined capability against source-backed reference answers. Inspect the model output, evidence, and grading rule behind each result.

How an evaluation works

  1. Scope the task

    Agree on questions, source coverage, and what counts as a correct answer.

  2. Build the proof set

    Preserve reference answers and provenance, including unsupported cases.

  3. Run and grade

    Retain model outputs and apply the agreed rules.

  4. Review the failures

    Inspect patterns and choose a focused next step.

What you receive

A baseline, evidence-backed failure examples, and the agreed export or report. Remediation and follow-up validation are separately scoped.

Custom evaluations are operator-led. We agree on price and turnaround before work starts.

A recorded example

Our SEC proof run contains 500 cases: 136 correct and 364 incorrect, a 27.2% baseline. This is our own task-specific evaluation, not an independent or universal model score.

Read the study and its limits →

Start with a concrete question.

Scope an evaluation