The measured record

Evidence before assurance.

BCOT Core does not claim to make a model omniscient. It changes which outputs are allowed to become commitments—and preserves the evidence for every decision.

Selective accuracy54.0%

of committed answers were correct

Up from a 28.2% baseline at 89.2% coverage.

252 of 587 wrong-answer commitments intercepted

TruthfulQA MC1 · external governance

Better commitments. No model retraining.

The same Llama-3.1-8B model moved from 28.2% baseline accuracy to 54.0% selective accuracy after the PAO boundary—while still answering 89.2% of questions.

Governed commitments54.0%

accuracy of answers that reached commitment

Coverage89.2%
What the third state contributed+30.3 pp

Observe lift

The lift from single-model governance to the full multi-architecture PAO run, using frozen operators.

252 / 587wrong-answer commitments intercepted

Keep the measures separate: 42.9% ΔHCR is the share of wrong commitments prevented. 54.0% is the accuracy of answers that did commit, at 89.2% coverage.

HaluEval · source-grounded summaries

AI that shows its receipts.

500 documents produced 1,000 evaluated summaries: 500 faithful and 500 fabricated. The analysis plan was locked and hashed before the run finished.

Frozen engine · pre-registered · zero post-lock changes

Fabricated summaries83%

never received a pass

57% contradicted with proof26% Observe · evidence resolution17% reached commitment
74.8%selective accuracy

On committed decisions at 70.3% coverage.

71.5–77.9%95% confidence interval

Reported around selective accuracy, not hidden behind a rounded headline.

84residual cases dissected

Every slip shared the same atomic-claim failure pattern.

Failure analysis

The fabrication lived in the binding between facts.

Break “Police in New York detained 60 protesters” into small enough claims and each fragment may appear somewhere in the source—even when the complete sentence is false because the event happened in Baltimore. Across two models, three configurations, and independent samples, every analyzed slip dissolved this way.

We publish the failure analysis. Ask your current vendor for theirs.

HALLMARK · 831 unseen citations

Coverage is a dial. Uncertainty is not hidden.

Real papers were mixed with fourteen fabrication types. The frozen governor ran once and returned one of three states: verified, contradicted, or cannot confirm.

Committed F11.0017% coverage
Strict operating point

144 verdicts · zero errors

Commit only when the evidence can sustain a definitive verdict.

0false accusations

Citation fabrication study

Fluency increased. So did fabrication.

Across 379 generated citations on three model architectures, 35.1% were fabricated. The instruction-tuned model produced the highest rate while also raising token confidence.

133 / 379fabricated citations

All 133 were intercepted when every citation was routed through retrieval-bounded verification using CrossRef and Semantic Scholar.

Gemma-7B
19.4%
Llama-3.1-8B base
31.1%
Mistral-7B Instruct
47.1%
Overall
35.1%

Detection is sample-scoped and retrieval-bounded. Retrieval coverage is the ceiling; this is not a universal “100% governance catch rate.”

TruthfulQA Regime C · substrate failure

What if every model learned the same wrong answer?

External governance needs a counter-signal. When six architectures agree on the same wrong answer, consensus can become the failure—not the cure.

Six-model persistence98.2%

332 of 338 consensus-wrong questions survived.

The pattern persisted across six open-weight architectures from three labs. A separate cross-lineage probe extended the presuppositional-trap behavior to eight model lineages; more inference did not surface the hidden premise.

persistent shared errordid not persist
The formal boundary

Inside the reachable regime, PAO can act on contradiction, missing evidence, or module disagreement. Outside it, the bottleneck is the training substrate itself. This is an existence-and-persistence result—not a claim about how often bad information appears in the world.

Evidence discipline

The record is part of the result.

A benchmark score is easy to polish after the fact. BCOT Core binds the plan, the frozen system, every verdict, and the report into a tamper-evident history.

  1. 01Pre-register

    Lock thresholds, metrics, and predictions.

  2. 02Freeze

    Preserve the exact engine and operators.

  3. 03Run

    Keep verdicts, evidence, confidence, and outcomes.

  4. 04Hash-chain

    Make later changes visible instead of deniable.

0post-lock HaluEval changes
2,290controlled early operator trials
3states recorded as first-class facts
How to read selective accuracy

Selective accuracy describes decisions that committed. It must always be read with coverage, which shows how often the system committed rather than routed to review.

What “all detected” means for citations

The 133 fabricated citations were detected because every citation was routed to retrieval-bounded verification. Detection remains limited by retrieval coverage.

What these results do not promise

Benchmark performance is not a universal guarantee. Deployed performance varies by domain, evidence quality, module configuration, and the cost you assign to review.

Bring your own commitments

See where three-state governance changes your risk.

In an Architecture Review, we map one real AI workflow to its commitment boundary, evidence sources, and measurable operating tradeoffs.

Explore the architecture

Architecture Review

Bring one commitment. Leave with its boundary.

Tell us a little about your architecture and the commitment you want to review.

We’ll use these details only to respond to your request. Privacy