Skip to main content
Use Cases ••13 min read•

Audit-Ready LLM Document Review Means Provable Coverage

Audit-ready LLM document review, decoded from AWS's Amazon Quick lease sweep: the Adjudicated Query pattern for provable coverage of every document.

Audit-ready LLM document review pairs generative models with deterministic sweeps and append-only records so every lease compliance finding can survive an audit.

The demo works. A compliance officer types a question into a chat window and gets an answer with counts attached: 10,111 lease-rule findings in breach, 689 ambiguous, 20 unreadable, every record accounted for. What AWS built for lease compliance in Amazon Quick is worth your attention, but not for the reason the demo suggests. Regulated document review with LLMs does not usually fail at the demo; it fails months later, when an auditor or opposing counsel asks which documents were checked, against which version of which statute, and by what method. This piece reads the Amazon Quick lease-compliance sweep as a build spec for audit-ready LLM document review: how the Adjudicated Query pattern manufactures provable coverage, what a full sweep costs, what an adjudication record must contain to count as evidence, and which parts of the design transfer to your own compliance workloads.

Stay in the loop.

Get the latest posts and exclusive content delivered to your inbox.

Join 10 readers. No spam. Unsubscribe in one click, anytime.

What the Amazon Quick Lease Sweep Actually Does

Watch the demo with an auditor's eyes and it holds up better than the usual LLM demo. The chat answer renders the completeness receipt as counts, and the dashboard drill-down behind it carries per-record rule versions, citations, and compared values. Which documents, which statute version, what method: the intro's three questions are answered on the screen. What survives is by whom, because the chat authenticates machine-to-machine; the records section returns to that gap. The subtler gap is durability. A chat answer is an assertion at a point in time; evidence has to survive paraphrase and personnel changes, and that is a property of a stored record, not a rendered answer.

A portfolio operator holds 50,000 apartment leases across multiple states, each governed by landlord-tenant statutes that change on the legislature's schedule rather than the operator's, and the build described in AWS's Adjudicated Query post pairs a conversational surface in Amazon Quick with a deterministic rules engine behind it.

Two design choices carry the whole pattern, both fencing the model away from anything official. First, the model never makes an official determination: Amazon Bedrock, embeddings for semantic clause search plus Claude for qualitative reads, powers only the exploratory path, while official sweeps are set-based comparisons against versioned rule rows, executed as parameterized SQL. Second, every sweep asserts a completeness receipt before anything persists. Read that way, the walkthrough's Texas numbers, 10,111 violations, 689 ambiguous findings, and 20 unreadable documents, are the output of an invariant: a durable record of what was checked, against what, and where the leftovers went. That record is what an auditor can actually inspect, and it is the component worth copying.

Provable Coverage vs. Extraction Accuracy

Generative AI document review becomes audit-ready only when the pipeline can prove which documents were checked, not just how accurately it read them.

Most generative AI document review today is trust-me extraction: the model reads a document, emits a judgment, and the number on the screen is the entire evidence package. Adjudicated review inverts the deliverable, making the product the durable record that a judgment happened, against what, and over exactly which documents.

These are separate failure surfaces. A pipeline can read clauses with high accuracy and still fail its audit because it cannot demonstrate which documents were checked against which version of the law. Benchmarks will not rescue you. The LegalBench paper measures how well models reason over legal tasks, which is genuinely useful for model selection, but a benchmark score says nothing about whether your Tuesday run covered the corpus. Audit-ready LLM document review is a property of the pipeline, not the model.

This is also why the AWS post is pointed about retrieval. Semantic search returns a ranked sample, and a ranked sample never knows what it excluded; no similarity threshold means all of them. Text-to-SQL looks more exact and is quietly worse, because a hallucinated predicate shrinks the population while the count still renders to the decimal. Provable coverage means something narrower and harder: every document in an enumerable population lands in exactly one disposition bucket, and the buckets sum to the population you claimed.

How the Adjudicated Query Pattern Works

The AWS post never states the pattern's most useful judgment: a fixed operation surface is itself an audit artifact. Because the model can only invoke a closed set of typed operations, an auditor can enumerate every capability it has and inspect the population logic inside each. That turns "the model never composes a wrong population" from a vendor promise into an externally checkable property.

The model's jobs are exactly two: translate the plain-language request into an invocation of one fixed typed operation, and narrate the structured result that returns. The reference build ships six tools: an exhaustive sweep, a rule-change simulation, an exploratory clause search, a single-finding drill-down, a rulebook listing, and a liveness check. The liveness check is transport evidence, not filler: proof the deterministic engine was reached and answering when the question was asked.

Rules are versioned data, not code. The engine knows generic comparisons (greater-or-equal, exists) and no branch naming a jurisdiction. A law change is a rulebook row edit carrying a version and citation, not a deployment. Every sweep asserts its receipt before committing:

compliant + in-breach + ambiguous + unreadable = scanned

A run that cannot account for its population never finishes, so no record is silently skipped. Findings are append-only, with no update or delete path anywhere in the code. Chat carries counts, the receipt, and a labeled sample; the dashboard carries the volume, so the model never summarizes away the guarantee.

One more mechanism deserves stealing. Even a narrator-only model reintroduces risk at the end of the chain, and AWS reports watching it happen: one narration stripped a caveat tag and invented a citation; another extrapolated a portfolio-wide range from twenty preview rows. The countermeasures are structural: bracketed caveat suffixes repeated at several payload levels, aggregates precomputed over every record, and mode labels, official versus exploratory, stamped wherever a number appears. The honest answer becomes the easy one.

What a Full Sweep Costs

Lease compliance automation has five cost drivers: portfolio size, rules per document, unit price, re-run cadence, and per-exception human review. The first three get all the attention. Re-run cadence usually dominates, and it multiplies the fifth.

Adjudication Volume per Full Sweep

Start with adjudication volume. A 50,000-lease portfolio checked against three rules per lease is 150,000 lease-rule pairs per full sweep. In the AWS reference build the sweep itself is set-based SQL, so its marginal cost is database compute and storage; the generative spend sits in extraction and the exploratory path, and the conversational layer bills per query under Amazon Quick pricing, so pull current rates instead of trusting any blog's number. To see the shape of the math, assume a placeholder half a cent per adjudication across extraction and querying combined. One full pass costs $750. That is the cheap part.

Re-Run Cadence Can Outrun Portfolio Size

Landlord-tenant statutes change on legislatures' schedules, tracked in the NCSL housing legislation database. If your answer to every amendment is a full-portfolio re-sweep, cost scales with change frequency, and cadence can outrun portfolio size entirely. A 10,000-lease portfolio re-swept monthly against three rules produces 360,000 adjudications a year; a 50,000-lease portfolio swept twice produces 300,000. The smaller portfolio costs more to keep provably current. The fix is scoping: wire each enacted amendment to a jurisdiction-scoped re-sweep covering only the leases under that state's law, with a low-frequency full sweep as a backstop. Statute monitoring is a cost lever as much as a compliance control.

The Exception Budget

Budget separately for the exception buckets. Ambiguous and unreadable documents route to humans, and that per-exception review cost recurs on every sweep: at an illustrative $2 per review, the walkthrough's 689 ambiguous findings add roughly $1,378 of recurring human labor to every pass, before re-runs multiply it.

The Record Behind Audit-Ready LLM Document Review

An LLM compliance audit trail stores rule versions, citations, dispositions, and timestamps so a regulator can inspect the evidence without re-running the model.

The Minimum Adjudication Record

When a regulator or a court inspects an AI-assisted review, they will not re-run your model. They will read your records. This is the LLM compliance audit trail in its minimal form, the record that turns model output into inspectable evidence:

FieldWhat it provesExample
document_idwhich record was checkedTX-LEASE-04471 (illustrative)
rule_id, version, citationwhich rule, which pinned law versionLATE_FEE_CAP v2026.01 (illustrative), TX Prop. Code ch. 92 subch. B
dispositionthe outcome, one of four bucketsin_breach
evidence_spanthe verbatim clause text"Late fee of 7% of monthly rent" (illustrative)
extracted vs required valuethe actual comparison performed7% against a 5% cap
model and config versionsreproducibility of the extractionmodel ID plus prompt and config hash
run_id and timestampwhich sweep, whensweep ID plus ISO timestamp (placeholder)
human overridewho changed what, and whyanalyst ID and rationale, or none
receipt referencepopulation completeness for the runcounts row reconciling to scanned

The 7 percent late fee, the 5 percent cap, and the Texas citation come from AWS's walkthrough. The lease ID, rule identifier, clause quote, and timestamp shown above are illustrative placeholders, not artifacts of the AWS sample. The transferable artifact is the field structure, not the Texas law.

Governance Anchors and the Identity Gap

Two governance anchors frame these AI governance records. The EU AI Act's Article 12 record-keeping duty requires high-risk systems to automatically log events over time, which is exactly what an append-only findings store provides. On AWS, Bedrock invocation logging captures model inputs and outputs with metadata at the account level, the plumbing behind the model-version field on any stack that uses Bedrock. And note the gap AWS itself flags: the chat integration authenticates machine-to-machine, so the token identifies the application, not the analyst who asked the question. If "by whom" must travel with the finding rather than living in a separate chat audit log, thread the end-user identity into the tool call and persist it on the sweep row.

Where These Pipelines Fail Their Audits

Four silent failure modes account for most audit collapses in this class of system. Each demands a named control in the design.

  1. Law drift. The sweep certifies the portfolio in March. An amendment takes effect in June, nobody re-runs, and the dashboard still says compliant, now against a law that no longer exists. The control is twofold: every record pins the law version it was judged under, and a statute monitor triggers a scoped re-sweep when that version is superseded. Without the pin, you cannot even measure your own staleness.
  2. Silent drops. Ingestion chokes on 340 scanned PDFs. They never reach analysis, no error surfaces, and the headline count looks exact because it is exact, for the wrong population. Generated queries make this worse, since a hallucinated predicate quietly shrinks the population while the number renders cleanly. The control is the completeness receipt asserted before persistence, with unreadable documents named as a bucket rather than dropped.
  3. Untracked overrides. An analyst flips a breach to compliant in a review UI. The rationale lives in an email thread, the analyst changes teams, and eighteen months later nobody can defend the change. The control is append-only records: an override is a new row carrying identity, rationale, and timestamp, never an in-place edit, with the original finding preserved beneath it.
  4. Hallucinated citations. The model cites a statute section that does not exist. Courts have sanctioned lawyers for filings built on ChatGPT-invented cases, the Mata v. Avianca episode being the standard example, and the tag-stripping failure described above shows the same risk surviving inside a narrator-only design. The control is structural. Citations live in the rulebook as versioned data, the model renders them but never sources them, and every rendered citation is diffed against the pinned statute text before it reaches a user, so a fabricated section number fails the check loudly instead of surviving to the audit.

What Transfers Beyond AWS

The portable core is four mechanisms, none of them AWS-shaped: pinning every judgment to a versioned rule and model, giving every document exactly one disposition, asserting a completeness receipt before persistence, and storing append-only evidence. The execution layer, Amazon Quick chat, QuickSight dashboards, Cognito token flows, the Lambda-hosted MCP server, Aurora storage, Bedrock models, is AWS plumbing. Swap it freely; the pattern does not care.

AWS's own fit test transfers just as well. Ask whether a missed record is a liability rather than an inconvenience, whether answers will be challenged later by someone who was not in the room, whether the governing logic is externally owned and changes on its own schedule, and whether the population is enumerable. KYC onboarding review passes all four: identity documents checked against sanctions and watchlist rules that regulators update without notice, over an enumerable applicant population. Insurance claims adjudication passes too, and AWS names it, along with export control screening, as candidate domains. Policy checks also fit, by our extension rather than AWS's list. A marketing content review fails the test, which is exactly the point of asking it.

A Minimal Build Spec for Any Compliance Sweep

Whatever stack you run, the build spec for provable coverage in AI compliance pipelines compresses to eight moves:

  1. Ingest against a manifest. Assign an ID and status to every document before any analysis. The population is defined here, and nowhere else.
  2. Make unreadable a disposition. Failed extraction lands in a named bucket, never in an exception log the pipeline swallows.
  3. Adjudicate against pinned rules. Rules are versioned rows with citations. No law reference lives in code.
  4. Assert the receipt before persisting. The four buckets must reconcile to scanned, or the run fails loudly and stores nothing official.
  5. Store the trail append-only. Every field from the table above, no update path, overrides as new rows.
  6. Log the model layer. Invocation logs with model and config identifiers are the reproducibility evidence behind the version fields; use Bedrock logging on AWS or the equivalent structured logs elsewhere.
  7. Wire change to re-sweep. Statute and rule updates trigger jurisdiction-scoped re-sweeps, with a periodic full pass as backstop.
  8. Keep the model at the edges. Translate questions, narrate results, rank exploratory samples. Never determine, never fix populations, never source citations.

Know when this is overkill. If the determination is genuine judgment, such as whether a clause is unconscionable, forcing it into a rules engine hides the subjectivity inside rule authorship and makes the output look exact when it is not. If the population is fuzzy, the receipt is precise about the wrong denominator. If plausible answers suffice and users can re-ask, plain RAG is the correct and cheaper answer. For enumerable populations under externally owned rules, though, the pattern's product is the thing compliance teams actually lack: a receipt that ages into evidence. The demo impresses; the record defends.

Stay in the loop.

Get the latest posts and exclusive content delivered to your inbox.

Join 10 readers. No spam. Unsubscribe in one click, anytime.

About the author

David Moreno

Applied AI Strategist

David helps teams put AI to work in real businesses. He writes teardowns of how companies actually deploy models: the architectures, the trade-offs, and the results that survive contact with the real world.

Related Posts