DreamersCertification Workbench

Evaluation framework

Accuracy is measured, not asserted: every loaded document has hand-verified ground truth, forming a known benchmark. Stakeholder tolerances — accuracy first, latency second, cost last — select the serving configuration. Post-award, the same framework runs on a stratified FLRA corpus with tolerances set jointly.

  1. Source of truth, known benchmark

    Every loaded document carries hand-verified ground truth: each field value, each expected absence, the page that proves it, every cited case. Extractions score against that truth — versioned and tracked run over run, so a change shows up as a number, not an anecdote.

  2. Stakeholder accuracy tolerance

    First in the optimization order. Stakeholders set the bar the benchmark must clear — including how an honest blank weighs against a confident guess. Tighten the bar and the selection changes.

  3. Latency tolerance

    Second in the order. How long each workflow can wait — interactive review versus overnight batch — decides which qualifying configurations remain. Cost breaks the remaining ties, last.

What every run records

A field the source does not state is scored correct only when left empty — an honest blank outranks a confident guess. Where abstaining and inferring disagree, the benchmark's policy decides which wins; that policy is itself a stakeholder choice, and the framework makes its effect visible.