Auditing LLM answers against the OCR text they cite
Blake McCarn
In early September I handed my document Q&A system a bad answer on purpose. The answer was "Your current insurance premium is $999." Nothing in the evidence supported it, and I made the verifier unavailable to simulate an outage. I ran it through strict mode, the mode that exists to refuse unsupported answers, and it came back unchanged.The system is Paperless Knowledge Graph. It sits on top of my Paperless-ngx archive and answers plain-language questions about scanned documents, with citations back to the source documents. I'd been calling its answers verified for months. The $999 test was one of fourteen probes in an accuracy audit I ran that week. Ten of them went after answers, citations and the eval suite, and all ten reproduced a defect.What I built afterward is a source audit. Before an LLM answer reaches the screen, every claim in it, and every citation behind those claims, gets checked against the original OCR text.
What "verified" used to mean
The old pipeline drafted an answer from retrieved chunks, then asked a second model call whether the answer was supported. That opinion fed a claim ledger and a numeric trust score, and the UI showed both. Each piece sounded reasonable on its own.The ledger accepted whatever references the model wrote. In one probe, a claim cited document 999999, evidence ID NOT-IN-PACK and the quote FABRICATED QUOTE, and the ledger counted it as supported. The trust score could also contradict its own ledger. I gave it a one-claim answer where the ledger said unsupported and a separate verifier said verified. The score came out at 0.945, rated high.The verifier also read only the first 6,000 characters of the answer and 8,000 characters of context, then applied its verdict to the whole answer. A fact past that cutoff was never checked, but it was graded.None of these failures needed a bad model. They needed a model that's wrong now and then, plus code that took its output at face value. That's the normal case for any LLM pipeline. So I didn't want a better verifier prompt. I wanted code to decide what counts as support, and to ask the model only the questions code can't answer.
Citations come from the archive, not the model
The first change was about where citations come from. A citation used to be text the model wrote, a document ID plus an evidence ID plus a quote. Any of those could be invented, and checking them afterward is a fuzzy matching problem with a lot of edge cases.Now the auditor never writes a quote. The app cuts the retrieved OCR into source windows, gives each one a handle, and sends them along with the claims. On the citation side, the auditor's only job is to point at handles. The app resolves each handle back to its document and original text. A handle that wasn't in the set the app sent gets rejected as unknown_or_unselected_span, and a claim with no valid reference fails before anyone asks whether it's true.Before any of that, the draft gets split into answer units, which are roughly sentences. Headings and "Premium:"-style labels don't become units of their own. They wait and get attached to the claim beneath them, because a label changes what the sentence under it means.Then a deterministic value check runs, whatever the auditor thinks. Every number and date in a claim has to appear in the passages that claim cites. If the answer says the premium is $1,284 and the cited window says $1,248, the claim fails with value_mismatch. Quantities have to match together with their units inside one contiguous passage, so "5 mg" in one quote and "90 kg" in another can't combine into "90 mg." Once every unit passes, the check runs again over the whole answer, since Markdown and list formatting can split a quantity across two units.A quote that matches proves provenance. It doesn't prove the claim. "Annual premium: $1,284" on a renewal offer and the same line on a bill are different facts. Telling them apart takes judgment, and that's where the model comes back in.
Six checks per claim
The auditor gets a batch of claims and the source windows. For each claim it starts by writing a short source basis, a plain description of what the original passages establish, before it judges anything. Then it fills in six checks, each one supported, not_established or contradicted.Take the claim "The annual premium is $1,284" and a renewal notice that contains that number. The subject check asks whose premium it is and for which policy, since one notice can cover several vehicles or policies. The predicate check asks whether $1,284 is really the annual premium, or the six-month installment, or last term's amount, or the discount on the next line. The record role check asks what kind of document this is and what its fields mean. A renewal offer states a price the insurer is proposing, and it doesn't show that anyone accepted it. Paperwork in general works this way. A signed direct-deposit form for a bank shows that I asked for the change. It doesn't show that the bank made it.Those three checks always apply. The response schema doesn't even let the auditor mark them not_applicable.The other three can be skipped when a claim doesn't need them. Conditions covers options and qualifiers. If the $1,284 assumes a paid-in-full discount, it isn't the premium unless that option was actually selected. Time covers what the dates mean. A notice from last March records what the insurer said last March, and nothing about today. Comparison covers words like "increased" or "latest," which only hold if every relevant record was retrieved and compared. A claim that makes a time or comparison assertion can't opt out of the matching check.Time needed the most rework. An early version treated any appearance of "latest" or "current" as a claim about the present world, so "the latest retrieved statement shows..." got suppressed even with every claim supported. Now the auditor tags each claim's time framing as a historical observation, a comparison across the retrieved documents, or a statement about the present, and only certain combinations of those tags are valid. A dated record can support the first two. It can't support the third, so a claim about the present needs evidence that settles the present, and a document date never does. If someone asks for their current premium and every claim in the answer is framed as what a dated record shows, the answer can still pass. The app appends its own line saying current status isn't established as of the evaluation date.Last, the auditor lists any assumption the claim needs that the sources don't state. For a supported claim, that list has to be empty. Here's an illustrative assessment for the renewal-offer case:Look at the last field. The checks say record role isn't established, and the auditor listed an assumption, but its verdict still says supported. That contradiction happens, so the parser doesn't take the verdict on trust. It collects every failed check and any assumptions as rejection reasons, and a "supported" verdict with reasons attached becomes unsupported:That's the reason for six checks instead of one yes or no. A yes is easy to give to a passage that shares the right words with the claim. Breaking the question apart makes the model commit, field by field, to each way the claim could be wrong, and when its own checks disagree with its verdict, the checks win. The failed checks also turn into specific repair instructions later, which beats a bare "unsupported."
Audit batches, and why every batch sees every source
Claims go to the auditor four at a time, with up to four batches running at once. The time budget is 60 seconds per round of batches. A twelve-claim answer is three batches in one round, so it gets 60 seconds. A twenty-claim answer is five batches and gets 120. Small batches keep each structured response short, and a failure stays contained to four claims.Every batch sees the full set of eligible source windows. An earlier version reranked and trimmed the evidence for each batch, and it failed in a predictable way. A claim could fail while the passage that supported it sat in the evidence pack, just not in the slice that batch got. Sending everything means every batch carries the same sources, which costs more input tokens. I'll pay that. A rejection should mean the sources don't support the claim, not that a budgeting heuristic guessed wrong.Each batch also gets the surrounding answer text, so a claim under a dated heading keeps its framing. The prompt is explicit that this context is for interpreting the claim and never counts as evidence for it.The retry rule took a while to get right. If a batch comes back malformed, with invalid JSON, a duplicate key, a renumbered unit ID or a handle that doesn't exist, it gets exactly one correction call with the error list and the same claims and sources. A negative judgment never gets a retry. If the auditor says a claim isn't supported, that's the result. Retrying judgments until one comes back positive is just sampling for a yes. The correction prompt says so directly. It tells the auditor to fix the structure, make a fresh source assessment, and not treat the correction as a request for a supported verdict.
One repair, then verified, partial or withheld
When the audit completes and some claims fail, a repair editor gets one attempt. It sees the question, the draft, the source windows and each claim's status with its rejection reasons, so it knows that one claim cited nothing, another stated a value the quote didn't contain, and a third failed on record role. It returns a list of one-line observations with no citations. Each observation has to name its own subject and date, because it'll be audited alone, without its neighbors. Then the revision gets a complete audit from scratch with a fresh ledger, and the app attaches citations only from references that validate.It's one repair and no more. Every repair costs another full audit, and a loop that keeps rewriting until something passes drifts toward whatever is easiest to cite rather than toward an answer to the question. If the editor hands back the same text, the loop stops there.After that, one of three things happens.The answer is verified when every unit was audited and supported and its time framing holds up. It's shown with document links built from the validated references.It's partial when every claim was audited and the failures are plain "unsupported" or "missing evidence." The app keeps only the supported claims and audits that smaller answer again as a new answer. Removing sentences can strand the ones left behind. "It went up the following year" means nothing once the sentence before it is gone, so the subset gets no credit from the first pass. If it passes, it's shown with a note saying some claims couldn't be verified and were omitted, and that the answer doesn't cover every part of the question. Partial answers are never cached. Conflicting sources rule this path out, and so do bad attributions, a value mismatch that only shows up across the whole answer, and invalid time or comparison framing. Those mean something in the answer is wrong, not just unsupported, and I don't want to trim around that.Everything else is withheld. That covers a timeout, an unreachable auditor, a batch still malformed after its correction, conflicting sources, or a subset that didn't pass. The user gets a fixed message: "I could not verify a complete answer from the retrieved source text." If the document index changes while the query is running, the answer is withheld too, with a message to retry once indexing finishes, since a clean audit against a changing set of sources doesn't mean much.The partial path wasn't there at first. Originally, if any claim still failed after the repair, the whole answer was replaced with the abstention. I ran into that on a history question where most claims passed and two failed the value check on dates. In one of them, the source table printed two-digit years, the answer used four-digit years, and the check refused to match them. Every other claim was verified, and all of it got thrown away. Withholding is the safe failure, but a system that withholds answers it has mostly verified is hard to keep using. The date problem got its own fix as well. When a short-year table and a full-year field both appear in the sources, the auditor is told to cite the full-year field.
The eval harness passed answers with no sources
The evaluation harness had its own version of the same problem. My original test set was six questions about my own archive, scored on required keywords and the system's own confidence. One probe built a fake response for each case out of the required keywords joined into a single string, zero sources, verification status not_run, a made-up trace, and a timeline event dated February 31, 2099. All six passed.That probe is in the test suite now with the assertion flipped:Scoring moved to a synthetic corpus with known document IDs, exact chunk text and an expected answer for each case. The scorer checks reference offsets and content digests against that fixture file, never against anything the model sent back. Confidence gets reported and has no effect on passing. A correct answer at confidence 0.01 passes. "The amount due is 152.00 USD" at 0.99 fails when the document says 125.00, and so does "125.00 EUR." Some cases expect the system to abstain, and an abstention with a guessed balance tacked onto the end fails.The six archive questions stayed as smoke checks. They still have to reach a final answer and cite real source IDs, but nobody ever labeled their source facts independently, so they don't count toward accuracy.The harness reports wrong answers and abstentions as separate numbers. A system that withholds everything never gives a wrong answer, and it's also useless. Neither number means much without the other.
What the source audit costs, and what it can't catch
Every audited answer costs at least one extra model call per four claims. A repair adds a call for the editor and a full second audit, and a partial answer adds one more audit for the subset. The audit policy has a version number that's part of the answer cache key, so when I tighten a check, answers certified under the old rules stop being served. It's at source-audit-v28 as of writing, which gives a sense of how often that has happened.Whether a passage means what the claim says is still a model's call. Code checks that references exist, that every unit was audited, that the structure is valid and that numbers and dates match. It can't check meaning, and every verification result carries a note saying semantic support is model-assessed.The audit can also only judge what retrieval found. If the right document never made it into the evidence pack, a claim about it fails as missing evidence, or an answer about the newest record ends up built around the newest record that happened to be retrieved. The audit makes that gap visible, because claims fail instead of passing quietly, but it can't close it. A history question needs every period a policy covered, and similarity search alone doesn't promise that.The value check is deliberately conservative, and it rejects some true claims. It used to reject a lot of table data. A statement table with "Charge USD" in the column header and 100 in the cell below doesn't contain the string "$100 USD," so a correct claim failed. The fix wasn't to loosen the match. The app now parses tables in the original OCR and binds a unit to a cell only through that table's own header or row label, never from some other excerpt.Signs are still strict. A credit printed as -$40 supports "the charge was reduced by $40," and the check still rejects that claim, because a check that ignored signs would also accept "a charge of $40" against the same credit. I'd rather reject that phrasing than accept an amount with the wrong sign.
{
"unit_id": "u3",
"source_basis": "A renewal offer proposing a 12-month premium of $1,284 for the term starting March 1. It does not record acceptance or payment.",
"checks": {
"subject": "supported",
"predicate": "supported",
"record_role": "not_established",
"conditions": "not_applicable",
"temporal": "supported",
"comparison": "not_applicable"
},
"unresolved_assumptions": ["The offered renewal was accepted and billed."],
"references": [{ "span_id": "s14" }],
"temporal_scope": "historical",
"temporal_assertion": "source_observation",
"comparison_scope": null,
"comparison_document_ids": [],
"status": "supported"
}
Python
reasons = ['semantic_' + facet for facet in FACETS
if checks[facet] in ('not_established', 'contradicted')]
if assumptions:
reasons.append('semantic_assumptions')
# ...
'status': 'unsupported' if row['status'] == 'supported' and reasons else row['status'],
Python
def test_all_six_old_keyword_only_forgeries_fail(self):
cases = json.loads((ROOT / "evals/canonical_questions.json").read_text())
for case in cases:
forged = {"answer": " ".join(case["required_terms"]), "confidence": 0.99, "sources": [],
"source_summary": {"trust_score": 0.99, "verification_status": "not_run"},
"trace": [{"made_up": True}], "timeline_events": [{"date": "2099-02-31"}]}
with self.subTest(case=case["id"]):
self.assertFalse(score_case(case, forged)["passed"])