Audit-Ready LLM Document Review Means Provable Coverage
Audit-ready LLM document review, decoded from AWS's Amazon Quick lease sweep: the Adjudicated Query pattern for provable coverage of every document.

In this article
- 1.What the Amazon Quick Lease Sweep Actually Does
- 2.Provable Coverage vs. Extraction Accuracy
- 3.How the Adjudicated Query Pattern Works
- 4.What a Full Sweep Costs
- 5.Adjudication Volume per Full Sweep
- 6.Re-Run Cadence Can Outrun Portfolio Size
- 7.The Exception Budget
- 8.The Record Behind Audit-Ready LLM Document Review
- 9.The Minimum Adjudication Record
- 10.Governance Anchors and the Identity Gap
- 11.Where These Pipelines Fail Their Audits
- 12.What Transfers Beyond AWS
- 13.A Minimal Build Spec for Any Compliance Sweep
The demo works. A compliance officer types a question into a chat window and gets an answer with counts attached: 10,111 lease-rule findings in breach, 689 ambiguous, 20 unreadable, every record accounted for. What AWS built for lease compliance in Amazon Quick is worth your attention, but not for the reason the demo suggests. Regulated document review with LLMs does not usually fail at the demo; it fails months later, when an auditor or opposing counsel asks which documents were checked, against which version of which statute, and by what method. This piece reads the Amazon Quick lease-compliance sweep as a build spec for audit-ready LLM document review: how the Adjudicated Query pattern manufactures provable coverage, what a full sweep costs, what an adjudication record must contain to count as evidence, and which parts of the design transfer to your own compliance workloads.
Stay in the loop.
Get the latest posts and exclusive content delivered to your inbox.
Join 10 readers. No spam. Unsubscribe in one click, anytime.
What the Amazon Quick Lease Sweep Actually Does
Watch the demo with an auditor's eyes and it holds up better than the usual LLM demo. The chat answer renders the completeness receipt as counts, and the dashboard drill-down behind it carries per-record rule versions, citations, and compared values. Which documents, which statute version, what method: the intro's three questions are answered on the screen. What survives is by whom, because the chat authenticates machine-to-machine; the records section returns to that gap. The subtler gap is durability. A chat answer is an assertion at a point in time; evidence has to survive paraphrase and personnel changes, and that is a property of a stored record, not a rendered answer.
A portfolio operator holds 50,000 apartment leases across multiple states, each governed by landlord-tenant statutes that change on the legislature's schedule rather than the operator's, and the build described in AWS's Adjudicated Query post pairs a conversational surface in Amazon Quick with a deterministic rules engine behind it.
Two design choices carry the whole pattern, both fencing the model away from anything official. First, the model never makes an official determination: Amazon Bedrock, embeddings for semantic clause search plus Claude for qualitative reads, powers only the exploratory path, while official sweeps are set-based comparisons against versioned rule rows, executed as parameterized SQL. Second, every sweep asserts a completeness receipt before anything persists. Read that way, the walkthrough's Texas numbers, 10,111 violations, 689 ambiguous findings, and 20 unreadable documents, are the output of an invariant: a durable record of what was checked, against what, and where the leftovers went. That record is what an auditor can actually inspect, and it is the component worth copying.
Provable Coverage vs. Extraction Accuracy

Most generative AI document review today is trust-me extraction: the model reads a document, emits a judgment, and the number on the screen is the entire evidence package. Adjudicated review inverts the deliverable, making the product the durable record that a judgment happened, against what, and over exactly which documents.
These are separate failure surfaces. A pipeline can read clauses with high accuracy and still fail its audit because it cannot demonstrate which documents were checked against which version of the law. Benchmarks will not rescue you. The LegalBench paper measures how well models reason over legal tasks, which is genuinely useful for model selection, but a benchmark score says nothing about whether your Tuesday run covered the corpus. Audit-ready LLM document review is a property of the pipeline, not the model.
This is also why the AWS post is pointed about retrieval. Semantic search returns a ranked sample, and a ranked sample never knows what it excluded; no similarity threshold means all of them. Text-to-SQL looks more exact and is quietly worse, because a hallucinated predicate shrinks the population while the count still renders to the decimal. Provable coverage means something narrower and harder: every document in an enumerable population lands in exactly one disposition bucket, and the buckets sum to the population you claimed.
How the Adjudicated Query Pattern Works
The AWS post never states the pattern's most useful judgment: a fixed operation surface is itself an audit artifact. Because the model can only invoke a closed set of typed operations, an auditor can enumerate every capability it has and inspect the population logic inside each. That turns "the model never composes a wrong population" from a vendor promise into an externally checkable property.
The model's jobs are exactly two: translate the plain-language request into an invocation of one fixed typed operation, and narrate the structured result that returns. The reference build ships six tools: an exhaustive sweep, a rule-change simulation, an exploratory clause search, a single-finding drill-down, a rulebook listing, and a liveness check. The liveness check is transport evidence, not filler: proof the deterministic engine was reached and answering when the question was asked.
Rules are versioned data, not code. The engine knows generic comparisons (greater-or-equal, exists) and no branch naming a jurisdiction. A law change is a rulebook row edit carrying a version and citation, not a deployment. Every sweep asserts its receipt before committing:
compliant + in-breach + ambiguous + unreadable = scanned
A run that cannot account for its population never finishes, so no record is silently skipped. Findings are append-only, with no update or delete path anywhere in the code. Chat carries counts, the receipt, and a labeled sample; the dashboard carries the volume, so the model never summarizes away the guarantee.
One more mechanism deserves stealing. Even a narrator-only model reintroduces risk at the end of the chain, and AWS reports watching it happen: one narration stripped a caveat tag and invented a citation; another extrapolated a portfolio-wide range from twenty preview rows. The countermeasures are structural: bracketed caveat suffixes repeated at several payload levels, aggregates precomputed over every record, and mode labels, official versus exploratory, stamped wherever a number appears. The honest answer becomes the easy one.
What a Full Sweep Costs
Lease compliance automation has five cost drivers: portfolio size, rules per document, unit price, re-run cadence, and per-exception human review. The first three get all the attention. Re-run cadence usually dominates, and it multiplies the fifth.
Adjudication Volume per Full Sweep
Start with adjudication volume. A 50,000-lease portfolio checked against three rules per lease is 150,000 lease-rule pairs per full sweep. In the AWS reference build the sweep itself is set-based SQL, so its marginal cost is database compute and storage; the generative spend sits in extraction and the exploratory path, and the conversational layer bills per query under Amazon Quick pricing, so pull current rates instead of trusting any blog's number. To see the shape of the math, assume a placeholder half a cent per adjudication across extraction and querying combined. One full pass costs $750. That is the cheap part.
Re-Run Cadence Can Outrun Portfolio Size
Landlord-tenant statutes change on legislatures' schedules, tracked in the NCSL housing legislation database. If your answer to every amendment is a full-portfolio re-sweep, cost scales with change frequency, and cadence can outrun portfolio size entirely. A 10,000-lease portfolio re-swept monthly against three rules produces 360,000 adjudications a year; a 50,000-lease portfolio swept twice produces 300,000. The smaller portfolio costs more to keep provably current. The fix is scoping: wire each enacted amendment to a jurisdiction-scoped re-sweep covering only the leases under that state's law, with a low-frequency full sweep as a backstop. Statute monitoring is a cost lever as much as a compliance control.
The Exception Budget
Budget separately for the exception buckets. Ambiguous and unreadable documents route to humans, and that per-exception review cost recurs on every sweep: at an illustrative $2 per review, the walkthrough's 689 ambiguous findings add roughly $1,378 of recurring human labor to every pass, before re-runs multiply it.
The Record Behind Audit-Ready LLM Document Review

The Minimum Adjudication Record
When a regulator or a court inspects an AI-assisted review, they will not re-run your model. They will read your records. This is the LLM compliance audit trail in its minimal form, the record that turns model output into inspectable evidence:
| Field | What it proves | Example |
|---|---|---|
| document_id | which record was checked | TX-LEASE-04471 (illustrative) |
| rule_id, version, citation | which rule, which pinned law version | LATE_FEE_CAP v2026.01 (illustrative), TX Prop. Code ch. 92 subch. B |
| disposition | the outcome, one of four buckets | in_breach |
| evidence_span | the verbatim clause text | "Late fee of 7% of monthly rent" (illustrative) |
| extracted vs required value | the actual comparison performed | 7% against a 5% cap |
| model and config versions | reproducibility of the extraction | model ID plus prompt and config hash |
| run_id and timestamp | which sweep, when | sweep ID plus ISO timestamp (placeholder) |
| human override | who changed what, and why | analyst ID and rationale, or none |
| receipt reference | population completeness for the run | counts row reconciling to scanned |
The 7 percent late fee, the 5 percent cap, and the Texas citation come from AWS's walkthrough. The lease ID, rule identifier, clause quote, and timestamp shown above are illustrative placeholders, not artifacts of the AWS sample. The transferable artifact is the field structure, not the Texas law.
Governance Anchors and the Identity Gap
Two governance anchors frame these AI governance records. The EU AI Act's Article 12 record-keeping duty requires high-risk systems to automatically log events over time, which is exactly what an append-only findings store provides. On AWS, Bedrock invocation logging captures model inputs and outputs with metadata at the account level, the plumbing behind the model-version field on any stack that uses Bedrock. And note the gap AWS itself flags: the chat integration authenticates machine-to-machine, so the token identifies the application, not the analyst who asked the question. If "by whom" must travel with the finding rather than living in a separate chat audit log, thread the end-user identity into the tool call and persist it on the sweep row.
Where These Pipelines Fail Their Audits
Four silent failure modes account for most audit collapses in this class of system. Each demands a named control in the design.
- Law drift. The sweep certifies the portfolio in March. An amendment takes effect in June, nobody re-runs, and the dashboard still says compliant, now against a law that no longer exists. The control is twofold: every record pins the law version it was judged under, and a statute monitor triggers a scoped re-sweep when that version is superseded. Without the pin, you cannot even measure your own staleness.
- Silent drops. Ingestion chokes on 340 scanned PDFs. They never reach analysis, no error surfaces, and the headline count looks exact because it is exact, for the wrong population. Generated queries make this worse, since a hallucinated predicate quietly shrinks the population while the number renders cleanly. The control is the completeness receipt asserted before persistence, with unreadable documents named as a bucket rather than dropped.
- Untracked overrides. An analyst flips a breach to compliant in a review UI. The rationale lives in an email thread, the analyst changes teams, and eighteen months later nobody can defend the change. The control is append-only records: an override is a new row carrying identity, rationale, and timestamp, never an in-place edit, with the original finding preserved beneath it.
- Hallucinated citations. The model cites a statute section that does not exist. Courts have sanctioned lawyers for filings built on ChatGPT-invented cases, the Mata v. Avianca episode being the standard example, and the tag-stripping failure described above shows the same risk surviving inside a narrator-only design. The control is structural. Citations live in the rulebook as versioned data, the model renders them but never sources them, and every rendered citation is diffed against the pinned statute text before it reaches a user, so a fabricated section number fails the check loudly instead of surviving to the audit.
What Transfers Beyond AWS
The portable core is four mechanisms, none of them AWS-shaped: pinning every judgment to a versioned rule and model, giving every document exactly one disposition, asserting a completeness receipt before persistence, and storing append-only evidence. The execution layer, Amazon Quick chat, QuickSight dashboards, Cognito token flows, the Lambda-hosted MCP server, Aurora storage, Bedrock models, is AWS plumbing. Swap it freely; the pattern does not care.
AWS's own fit test transfers just as well. Ask whether a missed record is a liability rather than an inconvenience, whether answers will be challenged later by someone who was not in the room, whether the governing logic is externally owned and changes on its own schedule, and whether the population is enumerable. KYC onboarding review passes all four: identity documents checked against sanctions and watchlist rules that regulators update without notice, over an enumerable applicant population. Insurance claims adjudication passes too, and AWS names it, along with export control screening, as candidate domains. Policy checks also fit, by our extension rather than AWS's list. A marketing content review fails the test, which is exactly the point of asking it.
A Minimal Build Spec for Any Compliance Sweep
Whatever stack you run, the build spec for provable coverage in AI compliance pipelines compresses to eight moves:
- Ingest against a manifest. Assign an ID and status to every document before any analysis. The population is defined here, and nowhere else.
- Make unreadable a disposition. Failed extraction lands in a named bucket, never in an exception log the pipeline swallows.
- Adjudicate against pinned rules. Rules are versioned rows with citations. No law reference lives in code.
- Assert the receipt before persisting. The four buckets must reconcile to scanned, or the run fails loudly and stores nothing official.
- Store the trail append-only. Every field from the table above, no update path, overrides as new rows.
- Log the model layer. Invocation logs with model and config identifiers are the reproducibility evidence behind the version fields; use Bedrock logging on AWS or the equivalent structured logs elsewhere.
- Wire change to re-sweep. Statute and rule updates trigger jurisdiction-scoped re-sweeps, with a periodic full pass as backstop.
- Keep the model at the edges. Translate questions, narrate results, rank exploratory samples. Never determine, never fix populations, never source citations.
Know when this is overkill. If the determination is genuine judgment, such as whether a clause is unconscionable, forcing it into a rules engine hides the subjectivity inside rule authorship and makes the output look exact when it is not. If the population is fuzzy, the receipt is precise about the wrong denominator. If plausible answers suffice and users can re-ask, plain RAG is the correct and cheaper answer. For enumerable populations under externally owned rules, though, the pattern's product is the thing compliance teams actually lack: a receipt that ages into evidence. The demo impresses; the record defends.
Stay in the loop.
Get the latest posts and exclusive content delivered to your inbox.
Join 10 readers. No spam. Unsubscribe in one click, anytime.
About the author
David Moreno
Applied AI Strategist
David helps teams put AI to work in real businesses. He writes teardowns of how companies actually deploy models: the architectures, the trade-offs, and the results that survive contact with the real world.
Related Posts
OpenAI Proaction Case Study Decoded, Claim by Claim
The OpenAI Proaction case study decoded claim by claim, from the 60% sales lift to the 75+ hours saved and what each measures before you copy the stack.
Is a Consumer AI Agent Worth It? $550 Saved, $64 Wasted
Wired's consumer AI agent trial audited. $550 saved, $64 wasted, with per-task economics, failure modes, security risks, and the adopt-versus-build call.
Devin GPT-6 Astra Self-Testing Moves Review to Evidence
Devin GPT-6 Astra self-testing shifts code review from reading diffs to auditing evidence. Here is the pattern, its failure modes, and audit criteria.


