Skip to content

Published · 8 min read

AI Contract Review: How it Works and What it Actually Catches

Explore how AI contract review works so you can identify risks, understand its limits, protect sensitive data, and know when to seek legal advice.

AI contract review is the use of large language models to read an agreement, extract its key terms, and flag clauses that may require additional attention, producing a first-pass assessment in seconds rather than hours.

The speed is real. In Better Call GPT, a 2024 study comparing large language models against lawyers on contract review, advanced models matched or exceeded human accuracy at identifying legal issues, at a fraction of the cost.

Accuracy is a different question, though. Large Legal Fictions, published in the Journal of Legal Analysis in 2024 by researchers at Stanford's RegLab and Yale, found that the same generation of models hallucinated on legal queries between 58% and 88% of the time, depending on the model.

Both findings hold at once, and that's the problem you need to solve before you sign anything an AI has reviewed.

This guide covers what AI contract review actually does, what the data show it catches, and how to decide when a tool just isn’t enough.

What is AI contract review?

People use the phrase for different things, and the difference determines who is responsible when the review is wrong:

1. Pasting a contract into a general-purpose chatbot: The model has no representation of what a complete contract should contain, no playbook of acceptable positions, and no way to rank severity.

2. A purpose-built review tool: Software is tuned for contract analysis, such as extracting clauses, comparing against a playbook, and flagging deviations. Still, there is no accountable human on the other end.

3. An AI-native law firm: Software does the first pass; a licensed attorney reviews it, exercises judgment, and signs off. The accountability sits with a person holding a bar license.

At its core, AI contract review means using language models for first-pass issue detection, clause extraction, deviation-from-playbook flagging, internal-consistency scrubbing, and redlining. None of the three produces a legal opinion.

How does AI contract review work?

Whether it's a standalone tool or the engine inside a firm's workflow, the mechanics are broadly the same.

Step What happens Where it breaks
1. Ingest and parse Converts PDF, Word, or scanned image into structured text Poor scans and unusual formatting corrupt downstream steps
2. Classify Identifies contract type, which selects the playbook and clause library Misclassification applies the wrong standard to everything after
3. Extract terms Pulls liability caps, indemnities, renewal windows, IP assignment, governing law Clause boundaries get drawn wrong, so the flag points at the wrong text
4. Compare to playbook Checks terms against acceptable ranges or fallback positions A generic playbook has no idea what's market for your deal size
5. Flag and redline Marks deviations, ambiguity, inconsistency; proposes language Everything gets flagged at equal weight
6. Attorney judgment A person applies leverage, business context, and negotiating strategy Skipped entirely in tool-only workflows

That last row carries more weight than most marketing copy suggests, and there's a reason. ContractEval, a 2025 benchmark for clause-level legal risk identification, found that most LLMs perform "at a level comparable to junior legal assistants" on this task. It also found open-source models generate "no related clause" responses even when relevant clauses are present in the document, which the authors attribute to "laziness" in reasoning or low confidence in extraction.

What does AI contract review catch well?

The case for AI here is real, and it’s narrower than it looks at first. It performs well on three things:

  • It flags risky clauses that are present in the text: If a liability cap, an auto-renewal, or an unusual indemnification carve-out is in the document, extraction models are good at finding it and flagging it.
  • It compresses the hours spent reading line by line: Thomson Reuters research finds lawyers spend 40% to 60% of their time drafting and reviewing contracts and other documents. A first pass that takes seconds removes most of the mechanical part of that work, which is the part that doesn't require a law degree.
  • It replaces the junior pass, not the attorney: A Better Call GPT study measured a 99.97% cost reduction against traditional review. That comparison runs against junior associate and legal-process-outsourcer review, not senior judgment.
Where AI contract review works the best

Read more: Benefits of AI in Law vs. Its Limitations: A Complete Guide

5 things AI contract review still misses

If the task is "find the obvious risk clauses fast," AI is now better than a distracted human doing the same job at 11 pm before a deadline. But the list of what it catches is still shorter than the list of what it misses.

1. The clause that isn't there

AI analyzes what's on the page. It is structurally bad at flagging what should be there and isn't, because absence produces no text to detect.

The cost of that gap is not hypothetical. CLAUSE, a benchmark built on more than 7,500 perturbed real-world contracts, opens with Perini Corp. v. Greate Bay Hotel & Casino, where an omitted consequential-damages waiver produced a $14.5 million liability, more than 20 times the underlying contract fee.

Across those contracts, the CLAUSE authors found that models "often miss subtle errors and struggle even more to justify them legally."

2. Internal consistency

Consistency asks whether what a contract says on page four matches what it says on page fifty-one, and checking a document against itself is the job you'd expect a machine to do well.

ContractScrub, a 2026 benchmark from Thomson Reuters Foundational Research and Imperial College London, tested frontier models on contracts hand-built by experienced lawyers containing misused defined terms, incorrect references, and inconsistent language.

The models performed surprisingly poorly. The best one found roughly three-quarters of the planted errors, and every model tested scored below 0.650 on the combined measure of catching real problems without flagging false ones. Also, where a term appeared capitalized as though it had been defined, but no definition existed anywhere in the document, average detection was 35%.

3. Hallucination

Large Legal Fictions found that the models showed "a tendency towards overconfidence, irrespective of their actual accuracy," and a "contra-factual bias"—assuming a premise in your question is true even when it's flatly wrong.

One caveat worth stating is that the study tested 2023-generation models. While the current models have improved since, the behavioral patterns, such as overconfidence and accepting your framing, have not been shown to disappear.

4. Market standard vs. deal-breaker

A model has no reliable way to know whether an indemnification cap is standard for your deal size and industry or wildly aggressive. It can tell you a clause exists, but it can't tell you whether to push back.

Also, flagging ten issues at equal weight is not the same as knowing which two are deal-breakers and which eight are boilerplate you can live with.

5. Confidentiality exposure

Every contract pasted into a public AI tool potentially exposes counterparty terms, deal economics, and information covered by an existing NDA. For a document with confidentiality obligations attached, that is a potential breach committed before the review even finishes.

When is an AI tool enough, and when do you need an attorney?

Not every contract carries the same stakes, so match the approach to what's actually on the table.

Situation Reasonable approach
Mutual NDA on standard terms AI first pass, usually fine to sign
Vendor agreement under $50k, standard SaaS terms AI first pass plus a quick attorney spot-check on flagged items
$250k+ MSA with custom indemnification or liability terms Attorney review required; AI accelerates but doesn't replace
Employment agreement with IP assignment or non-compete Attorney review required; jurisdictional variance is high-stakes
Anything with confidential counterparty terms Never paste into a public tool; use a confidential, attorney-reviewed workflow

AI handles the top two rows well enough that paying for attorney review is often unnecessary. The failure modes are concentrated exactly where the money is.

Bloomberg Law surveyed in-house and law firm attorneys in 2025 and found that 63% had used AI for work in some form. What they used it for clusters at one end of the job:

  • 39% for drafting memos, emails, and correspondence
  • 37% for summarizing case law
  • 30% for reviewing legal documents

For contract negotiation and drafting legal agreements, Bloomberg Law's analysis found that the vast majority of respondents hadn't adopted AI tools at all, and attributed the gap to concerns about data security, compliance, and professional responsibility.

Read More: Contract Review & the AI-Native Model

Most firms bolted a drafting assistant onto an unchanged workflow, producing an efficiency gain that stops at the firm.

General Legal is an AI-native law firm for companies signing more contracts than a founder can reasonably read. Its contract AI, Sentinel, produces redlines, annotations, and an issues list in roughly ten seconds. An attorney then spends 15 to 30 minutes on the work that used to take hours (market judgment, risk severity, negotiation strategy, and the omissions a model can't see). Work arrives and returns through a private Slack channel, with a median first turn of three hours or less.

The attorneys doing that second pass come from firms including Fenwick, Cooley, Latham & Watkins, and Morgan Lewis. Pricing is quoted upfront as a flat fee.

The result is a structure that answers the specific failure modes rather than talking around them:

  • Omissions and completeness are an attorney's job, because they are not detectable in the text.
  • Severity ranking comes from someone who knows your leverage and your deal, not from a flat list of flags.
  • Confidentiality holds because the contract never enters a public tool.
  • The price is published, so the speed gain reaches you instead of stopping at the firm.

Have a contract blocking a deal right now? You can create a free account with no commitment or upfront fees and send the document for a quote, or book a 10-minute working session to discuss how your contracts move today.

Key takeaways
  • AI contract review excels at finding risky clauses present in text and compressing review time, but performs at junior legal assistant level on clause-level risk identification.
  • Advanced models hallucinate on legal queries 58% to 88% of the time according to Stanford and Yale research, even while matching human accuracy on specific contract review tasks.
  • AI systematically misses absent clauses, struggles with internal consistency (detecting only 35% of undefined capitalized terms), and cannot distinguish market-standard terms from deal-breakers.
  • Pasting contracts into public AI tools creates confidentiality breach risks; high-stakes agreements with custom terms or significant dollar exposure require attorney review.
  • AI-native firms combine machine first-pass review with licensed attorney judgment to address AI's structural limitations while maintaining speed and cost advantages.
TL;DR
What AI catches wellIdentifies risky clauses present in text, compresses review time from hours to seconds, and replaces junior-level mechanical reading tasks
Accuracy paradoxModels match human accuracy on contract issue identification but hallucinate 58-88% on legal queries, with both findings holding simultaneously
Missing clauses problemAI cannot flag what isn't there—omitted clauses have produced multi-million dollar liabilities in real cases
Internal consistency gapsBest models catch only three-quarters of planted errors and detect just 35% of capitalized but undefined terms
Confidentiality exposurePasting contracts into public AI tools potentially breaches existing NDAs and exposes deal economics before review completes
When attorney review requiredAgreements over $250k, those with custom liability terms, employment contracts with IP assignment, and anything with confidential counterparty terms need licensed attorney review
AI-native firm modelCombines AI first-pass analysis with attorney judgment on omissions, severity ranking, and strategy while maintaining confidentiality and speed
FAQs

Can I use AI to review contracts?

Yes, and for routine, low-stakes documents, it's often the right first move. For anything with real dollar exposure, custom terms, or negotiation ahead of it, that pass needs an attorney checking what the model missed.

Can ChatGPT review contracts?

It can produce a first-pass read with real structural limits. However, a general-purpose model has no fixed representation of what a complete contract should contain, no curated playbook of acceptable terms, and no reliable way to rank severity.

How accurate is AI contract review?

It depends on the task. On identifying legal issues present in the text, research has found advanced models matching or exceeding human accuracy. On clause-level risk identification, studies place most models "at a level comparable to junior legal assistants." On legal questions generally, hallucination rates in published research have run from 58% to 88%, depending on the model.

Is a contract reviewed by AI legally valid?

Yes. Under the E-SIGN Act of 2000 and the Uniform Electronic Transactions Act, a contract can't be denied legal effect merely because an electronic agent was involved, provided the action is attributable to the party meant to be bound. UETA §14 goes further, allowing formation through electronic agents even where no individual reviewed the specific terms.