The first two pieces framed the problem and the tool. The problem: an LLM is optimized to please, not to tell the truth โ it returns the most agreeable answer, and you can read it off the variance of repeated answers (confident-but-wrong = entrenched bias). The tool, precorrect: instead of filtering after the fact, it probes the model, finds which distortions it slides toward, and injects the correction upstream โ measuring each answer against an answer key, the standard of what's true, backed by solid sources. One promise was left to prove: how far can you really move a model toward the truth โ and what's the lever? Here are the numbers, on one-by-one verified keys โ including the part I didn't expect.
The setup: two engines
For TeoCentro I built an internal, vertical engine for bias discovery, hand-tuned on theology: doctrinal lenses refined document by document, and an answer key distilled from the corpus โ the standard of what, in that domain, is correct. It's the deep system, the work of years: the moat. Discovery runs once over the whole corpus, then refines continuously.
Then I took the same mechanism and stripped the domain knowledge out of it: that's the generic engine, precorrect โ discovery, correction and verification, but empty. The lenses and the key aren't in the code; they're what you feed it, and that's where the power sits. The generic engine ships bare; the internal one carries the moat.
The practical question: how far can you really move a model toward the truth, and what's the lever? Everything below is measured on one-by-one verified keys, never on an LLM-asserted key. The mechanism is generic (you carry it anywhere); the answer key โ the standard of what's true in your domain, backed by sources โ is what makes the difference.
The part that didn't work โ and why that's the honest finding
The obvious next step: take those discovered corrections, inject them into generation, and watch the articles improve. I ran that A/B (no rules / source rules / truth rules), many repetitions, scored for doctrinal accuracy. Injecting rules did not beat the no-rules baseline. On a capable model and well-covered topics, the bare model was already near the ceiling; piling rules on top added variance โ sometimes better, often worse.
This isn't a failure; it's the sharper version of the lesson. Anthropic's interpretability work describes persona vectors โ traits like sycophancy or honesty encoded as directions in the model's activations, steerable by prompts and examples. Crude rule-injection doesn't cleanly steer that vector; it jostles it โ which is why the variance explodes. What works is re-aiming the persona toward a truth-seeking frame (the grounding), not accumulating rules on top of it. Correction is a matter of direction, not volume.
Stuck โ and confident
So how stuck is the model? On one-by-one verified keys, asking each question five times, the errors are stable, not noise: between 16% and 36% of questions (depending on the model) it gets stably wrong โ at least four times in five, the same wrong answer. It isn't unsure; it's confidently, repeatably wrong.
And it isn't a matter of size: the most capable model isn't wrong less often. If anything โ as we'll see โ it resists more. Capability moves the error rate; it doesn't dissolve the entrenchment.
How far you can move it
Then I measured it properly: on the cases where the model is reliably wrong (keys checked one by one), against the bare 0%:
| intervention | flips wrong โ right |
|---|---|
| a warning ("think carefully, don't assume the popular view") | ~17% |
| the raw source citation | ~18% |
| the focused correct reading (the answer key) | ~50% |
| the key + stacked verified sources | ~57% |
| the key + "reason step-by-step against the popular view" | ~16% โ collapses |
Two clear things. The right, focused fact is the lever: a warning or a raw citation move it little (~โ ); it's the correct reading that flips half the cases. And โ the fine point, counterintuitive โ it's not volume, it's deliberation: stacking more verified sources holds the line (it even helps), but inviting the model to reason step-by-step against the popular view makes it collapse โ reasoning hands it more to rationalize its way back to the error (the persona-vector jostling). It's direction, not deliberation. A pointer is not a grounding; the right fact is โ pile on sources if you like, but don't invite it to "think it over."
And it's not about talking to it: I tried moving it with a "you are a rigorous scholar" frame, even an identity frame backed by primary sources โ none beat a generic warning (~โ ). Only the verified fact moves it (~50%). It's the injected datum that counts, not the pep talk.
And the more capable the model, the more it resists: the same correct reading saves ~60% on Haiku, ~60% on Sonnet, but only ~31% on Opus (the biggest, most capable). Handed the right answer bare, the largest model still holds the wrong one about two times in three โ capability doesn't dissolve the entrenchment, it hardens it. Not that it's immovable: with stacked evidence it climbs toward half (~50%), but it needs more proof โ and inviting it to reason sinks it. (A pilot on one-by-one verified keys, ~125 cases across three models; the shape is sharp and consistent.) precorrect ships the mechanism empty; the right, focused grounding you bring is the whole game.
In the whole picture, these percentages are the entrenched core: they measure only the cases where the model is stuck. Across the full set of hard expert questions โ central, contested points of doctrine, not trivia โ the model already gets most right on its own; focused grounding recovers half of what remains, lifting overall accuracy toward ~90%. And the residue that resists even with the proof in front of it is not waste: it's the map of the truths LLMs won't admit โ a datum in its own right.
How deep it digs โ and the bedrock
That residue isn't a fixed wall. On the ~29 hardest cases โ where even the top model holds the wrong answer with the fact in front of it โ I ran a second pass: progressively deeper grounding (the exact primary source; then multiple converging sources and the verbatim original text). It digs: ~19 of 29 give way under progressive targeted evidence. But ~10 of 29 are true bedrock โ they resist even the verbatim original text and multiple sources together (mostly very specific specialist claims). Targeted grounding digs deep; but a core remains โ and honesty is to name it, not to pass it off as solved.
The gate exploits exactly this: skip where the model is already right, intervene where it's confidently wrong.
the model โโโโโโโโ ~74% already right & stable โโโโโโโโ โ โโ ~26% confidently wrong โโ
(here it knows) (here it doesn't)
precorrect โโโโโโ ~74% SKIP โ gate, ~ยพ saved โโโโโโ โ โโโ ~26% CORRECT โ inject grounding โโโ
(already right โ leave it) (confidently wrong โ re-ground it)
The tool spends its effort exactly where the model fails and stays out of the way where it doesn't.
It doesn't just answer wrong โ it stops early
Here's the tell nobody connects. The same economy that makes the model give the agreeable answer also makes it stop early: it has internalized a halting rule โ good enough, stop โ at the inferred-acceptable point, not the actually-finished one. The confident stable error and the half-finished task are one thing seen from two sides. Re-grounding doesn't only change what it says; it changes how far it's willing to go. Fix the standard it answers to, and it both gets more right and digs deeper.
How you measure any of this
One note on method, because it is the method. You must decide who grades the answer โ and you can't let the model grade itself: it shares its own blind spots, so it ratifies its confident errors instead of catching them (the documented self-evaluation inflation). Hand the judge the answer key and have it check alignment โ not "is this true?" but "does this match the known-correct answer?". Graded that way, open-ended replies can be scored without collapsing them to a coin-flip. And the judge must be anchored to real sources: on a control sample, such a verifier agrees with the known keys on about nine cases in ten, at high confidence. The external key stops being an afterthought and becomes the instrument โ like tuning a guitar before you play.
Where this goes โ a benchmark
Everything here came from one probe battery on a few models. Point it at every model as it ships โ by domain, repeated for variance, graded against an external answer key โ and you get something the field is short on: a benchmark for specialized, verifiable knowledge, the kind where the popular answer and the true one come apart. Math and code have their leaderboards; doctrine, halakhah, textual criticism โ where even frontier models are confidently wrong โ don't yet. A public leaderboard, methodology open, the battery itself held back so it can't be trained against: the same answer-key discipline that runs through everything above, turned outward โ the public face of what precorrect's evaluate() does privately against your corpus.
What it means
Two things, and they point the same way:
- The mechanism is generic and reproducible. The same engine works across groundings; the code isn't the moat.
- The moat is the answer key. What turns a fluent consensus-machine into a truth-teller is an external standard of what's true โ and, on the hard questions, that standard can't be derived from the model or from popularity. Someone has to supply it.
The first article gave this a name. An LLM is structurally anthrลpareskos (แผฮฝฮธฯฯฯฮฌฯฮตฯฮบฮฟฯ) โ a "people-pleaser," one who acts to be liked rather than to be right. An external truth-standard is what turns it, on the hard questions, toward philalฤthฤs (ฯฮนฮปฮฑฮปฮฎฮธฮทฯ) โ a "truth-lover": not by its nature, but by what you hold it to. Alignment to truth isn't something a model gives you for free. It's something you plant.
The tool: precorrect on GitHub. The problem it solves, and why: article 1 ยท article 2.
The series on truthful AI: 1. The problem ยท 2. The tool ยท 3. The proof (this).