The unit of evaluation
A generated answer, citation or action is a candidate until it crosses a commitment boundary. Our evaluations keep model quality and governance quality separate: model accuracy describes the candidates; selective accuracy, coverage and dispositions describe what governance allowed to become real.
The run is frozen before the result
- Define the candidate set, modules, thresholds and outcome measures before the production run.
- Freeze the exact engine and evidence operators used for evaluation.
- Record every module verdict, disposition, reason and evidence reference.
- Preserve corrections as additions so the original record remains reconstructable.
Measures reported together
- Selective accuracy: accuracy among candidates that committed.
- Coverage: the share of candidates that received a committed verdict.
- Wrong commitments intercepted: incorrect candidates that did not reach a user.
- Resolution cost: correct candidates routed through unnecessary review.
- Residual errors: wrong candidates that still committed, retained rather than hidden.
Public benchmark sources
The evaluations use publicly described benchmarks so readers can inspect the underlying task. The linked sources define the benchmark; BCOT Core's reported figures describe documented governance runs on selected public sets.
Limits on the claim
BCOT Core is not a truth oracle. Results depend on the modules, evidence and authority available to the boundary. Benchmark performance does not guarantee performance in a different domain, and holding a commitment for resolution has a measurable cost. We report both the protection and the cost.