On committed decisions at 70.3% coverage.
The measured record
Evidence before assurance.
BCOT Core does not claim to make a model omniscient. It changes which outputs are allowed to become commitments—and preserves the evidence for every decision.
of committed answers were correct
Up from a 28.2% baseline at 89.2% coverage.
TruthfulQA MC1 · external governance
Better commitments. No model retraining.
The same Llama-3.1-8B model moved from 28.2% baseline accuracy to 54.0% selective accuracy after the PAO boundary—while still answering 89.2% of questions.
accuracy of answers that reached commitment
Observe lift
The lift from single-model governance to the full multi-architecture PAO run, using frozen operators.
Keep the measures separate: 42.9% ΔHCR is the share of wrong commitments prevented. 54.0% is the accuracy of answers that did commit, at 89.2% coverage.
HaluEval · source-grounded summaries
AI that shows its receipts.
500 documents produced 1,000 evaluated summaries: 500 faithful and 500 fabricated. The analysis plan was locked and hashed before the run finished.
never received a pass
Reported around selective accuracy, not hidden behind a rounded headline.
Every slip shared the same atomic-claim failure pattern.
The fabrication lived in the binding between facts.
Break “Police in New York detained 60 protesters” into small enough claims and each fragment may appear somewhere in the source—even when the complete sentence is false because the event happened in Baltimore. Across two models, three configurations, and independent samples, every analyzed slip dissolved this way.
We publish the failure analysis. Ask your current vendor for theirs.HALLMARK · 831 unseen citations
Coverage is a dial. Uncertainty is not hidden.
Real papers were mixed with fourteen fabrication types. The frozen governor ran once and returned one of three states: verified, contradicted, or cannot confirm.
144 verdicts · zero errors
Commit only when the evidence can sustain a definitive verdict.
Citation fabrication study
Fluency increased. So did fabrication.
Across 379 generated citations on three model architectures, 35.1% were fabricated. The instruction-tuned model produced the highest rate while also raising token confidence.
All 133 were intercepted when every citation was routed through retrieval-bounded verification using CrossRef and Semantic Scholar.
Detection is sample-scoped and retrieval-bounded. Retrieval coverage is the ceiling; this is not a universal “100% governance catch rate.”
TruthfulQA Regime C · substrate failure
What if every model learned the same wrong answer?
External governance needs a counter-signal. When six architectures agree on the same wrong answer, consensus can become the failure—not the cure.
332 of 338 consensus-wrong questions survived.
The pattern persisted across six open-weight architectures from three labs. A separate cross-lineage probe extended the presuppositional-trap behavior to eight model lineages; more inference did not surface the hidden premise.
Inside the reachable regime, PAO can act on contradiction, missing evidence, or module disagreement. Outside it, the bottleneck is the training substrate itself. This is an existence-and-persistence result—not a claim about how often bad information appears in the world.
Evidence discipline
The record is part of the result.
A benchmark score is easy to polish after the fact. BCOT Core binds the plan, the frozen system, every verdict, and the report into a tamper-evident history.
- 01Pre-register
Lock thresholds, metrics, and predictions.
- 02Freeze
Preserve the exact engine and operators.
- 03Run
Keep verdicts, evidence, confidence, and outcomes.
- 04Hash-chain
Make later changes visible instead of deniable.
How to read selective accuracy
Selective accuracy describes decisions that committed. It must always be read with coverage, which shows how often the system committed rather than routed to review.
What “all detected” means for citations
The 133 fabricated citations were detected because every citation was routed to retrieval-bounded verification. Detection remains limited by retrieval coverage.
What these results do not promise
Benchmark performance is not a universal guarantee. Deployed performance varies by domain, evidence quality, module configuration, and the cost you assign to review.
Bring your own commitments
See where three-state governance changes your risk.
In an Architecture Review, we map one real AI workflow to its commitment boundary, evidence sources, and measurable operating tradeoffs.