There’s a number that turns up almost every time somebody tells you AI can review your contracts. 94% accurate, against 85% for human lawyers. 26 seconds, against 92 minutes.

It’s been quoted in law journals, in board decks, and in a lot of sales pitches. So I went and read the actual study. It’s a 2018 whitepaper from a contract-review company called LawGeex, testing its own product: 20 lawyers, four hours, 5 non-disclosure agreements pulled from the old Enron document set. It was never peer reviewed and never replicated in the eight years since, and the study page is no longer on the company’s own website.

And there’s a sentence in it I don’t think many people have gotten to.

Two things first. I’m not a lawyer, and nothing here is legal advice. What I am is an engineer who’s spent 30 years reading what suppliers will and won’t stand behind. Second, every number below comes with its source, including who paid for it.

The sentence in the study

Verbatim from their own report: “The scoring of contract reviews by participants factored in the best answers of all 21 participants (including the LawGeex AI) to create ‘model answers.’ This formed the benchmark for scoring the answers.”

Read it twice. The system being tested helped write the answer key it was graded against.

Once you’re looking, more of it comes apart. The “top corporate lawyers” in the write-up were recruited partly through Upwork, the freelance hiring site, and by LawGeex’s own participant list the group included a contracts administrator. The documents were 5 NDAs, which Ken Adams, who wrote A Manual of Style for Contract Drafting, calls the cockroach of the contract world. Meaning the most standardized thing you could possibly test on. Adams has a longer argument about whether NDAs should count as a test of contract review at all, and that’s a piece for another day.

Being fair about it

It was a real experiment. They published their method. That scoring sentence sits in their own report in plain sight, and they’re the ones who put it there. Nobody lied about anything.

What happened next is that eight years of people repeating a number stripped away every one of those qualifications until all that was left was 94 beats 85. That’s a citation problem, and citation problems are extremely common.

Why the pressure is real

Corporate legal departments are running leaner than they have in years. Total legal spend has fallen to 0.43% of company revenue, a six-year low, at a median of 3 lawyers per $1B of revenue (ACC and Major, Lindsey & Africa, 576 departments, June 2026). 81% of departments report rising matter volumes while 55% have flat or shrinking budgets (Thomson Reuters Legal Department Operations Index, 2025). And a survey run with Harvard Law School’s Center on the Legal Profession found contracting teams spend over 40% of their time and budget on unchallenging, low-complexity contracts (EY Law and Harvard, 2021, so pre-GenAI). If that’s your week, a machine that reads the boring ones isn’t a luxury.

What these tools actually do

You feed in a contract. The system pulls out the clauses, matches them against a playbook of what your company will and won’t accept, and flags the ones that don’t match. Some go further and suggest replacement wording, which is the part called redlining: marking up a contract with proposed changes. Then a human decides which flags actually matter.

Where AI genuinely wins

Last year eight law firms ran a proper benchmark, and the lawyers taking part weren’t told they were being measured against software (Vals Legal AI Report, February 2026). On answering questions about a document, the best tool scored 94.8 against the lawyers’ 70.1. On summarization, 77.2 against 50.3. On structured data extraction, 75.1 against 71.1. And on one specific extraction job, pulling 40 fields out of a credit agreement, the humans lost badly: 61.7% for the best tool against 39.2% for the lawyer group. If the task is find me every termination clause in these 200 agreements, the machine is better than you at it. Genuinely.

One disclosure, since I’m holding everyone else to it: Vals discloses a customer relationship with at least one participating firm, and the vendors chose which tasks to enter. Best independent benchmark that exists, not a perfect one.

Then they tested redlining

The lawyers scored 79.7. The best AI tool scored 65.

That was one of only two tasks out of seven where humans beat every single tool in the study. The evaluators found the tools “struggled to decipher the redlines from the original text,” and on amendments with several requirements they simply inserted standard text as a new clause instead of making the more nuanced changes the contract needed. So the thing most people bought this software to do is the thing it’s worst at.

The clause with no heading

One example from that benchmark stuck with me.

A contract contained a most-favoured-nation clause, which is a promise that you’ll get terms at least as good as anybody else gets. Except it wasn’t labelled that way. It just said the company would provide access to its personnel “no less favorable than what it provides any other customer.” The lawyers found it. One tool found it. One tool confidently returned something irrelevant. The rest reported that there was nothing there.

That’s the shape of the whole problem. The clause that costs you money is almost never the one with a helpful heading on it. The same benchmark ran a warranty-disclaimer test where the right answer was three separate clauses. Every tool found at least one. None found all three.

Two randomized trials

One benchmark is one benchmark, so look at the controlled trials, where people get randomly assigned to use the tool or not. There are two good ones, both out of Minnesota.

The first, 60 participants with blind grading, found no statistically significant quality improvement across any of its four tasks, though people were meaningfully faster, somewhere in the 12% to 32% range (Choi, Monahan and Schwarcz, 109 Minn. L. Rev. 147, 2024). The second, 137 participants across Minnesota and Michigan, did find real quality gains (Schwarcz et al., 2026).

Look at where, though. The gains landed on litigation work. On the one transactional task, drafting a short NDA, the researchers wrote that the gains “do not appear to extend to the one transactionally oriented task we evaluate, which involved drafting a short contract.” Neither tool improved any of the five quality attributes they measured, and neither meaningfully changed how long it took. Two independent randomized trials, and contract work is where the help runs out.

One more number, with a caveat attached

Stanford researchers ran a peer-reviewed, preregistered study of the big commercial legal AI products, against marketing that promised hallucination-free output (Magesh et al., Journal of Empirical Legal Studies 22(2), 2025). Lexis+ AI hallucinated on 17% of queries and came out accurate on 65%. Westlaw’s AI-Assisted Research hallucinated on roughly a third. Thomson Reuters’ Ask Practical Law hallucinated on 17% and was accurate on 19%.

The caveat matters and I’m not burying it: that study measured legal research, not contract review, and the authors say evaluating contract analysis is still an open challenge. What makes it relevant here is that Ask Practical Law is the transactional product, the one people reach for on deal work, and it landed at 19%.

Who owns the mistake

Say you use one of these, it misses a clause, and it costs your company real money. Who’s responsible? That answer has been sitting in the contract you agreed to when you bought the software. I read those too.

Spellbook, a contract drafting and review tool, states at §8.4(b)(i), in capitals: “SPELLBOOK DOES NOT REPRESENT OR WARRANT THAT THE SPELLBOOK AI PLATFORM WILL PRODUCE ACCURATE OR RELEVANT CONTENT FOR THE CUSTOMER.” Harvey, sold to large firms, says at §1.5, bolded in the original: “The Service is a research tool, and its Output is not legal advice.” LexisNexis, whose AI expressly generates contract clauses, writes at §12.11: “You assume all risks and liabilities in relying on the Online Services.”

Then there are the caps. What you can recover is typically limited to the fees you paid over the preceding 12 months. Harvey caps its liability for wrong output at $250,000 and caps a data breach separately at $500,000. Twice as much protection for leaking your document as for getting it wrong. That’s not a scandal. It’s a choice, written down, and you agreed to it.

Nobody has been to court yet

I went looking for the lawsuit, the case where an AI-reviewed contract went wrong and somebody sued. Court records, the legal press, the trade databases. As far as I can find it doesn’t exist: no adjudicated case anywhere from 2024 through 2026 over an AI-drafted or AI-reviewed contract, and no suit against one of these vendors over bad contract output.

Which sounds reassuring right up until you work out why. The risk got assigned before anything went wrong, in a terms-of-service page rather than a courtroom. And deal work is private, so a bad clause doesn’t make the news. It sits in a drawer until somebody tries to enforce it years later.

What we do have is a regulator and a tribunal. The FTC finalized an order against DoNotPay in February 2025, $193,000 in monetary relief, over an AI service whose advertised documents included NDAs, non-competes and residential leases. The complaint states that “DoNotPay employees have not tested the quality and accuracy of the legal documents,” and that subscribers complained the documents “were not fit for use.” In Moffatt v. Air Canada, 2024 BCCRT 149, a CAD $812.02 award over a chatbot that gave a passenger the wrong refund policy, the airline argued the chatbot was “a separate legal entity that is responsible for its own actions.” The tribunal called that “a remarkable submission.”

The insurers noticed

A survey of 13 professional liability insurers who between them cover more than 80% of Am Law 200 firms found 7 of them reporting a rise in AI-related claims, the first meaningful increase in five years (EPIC Law Firm Group, 16th Annual LPL Claims Survey, May 2026). The practice area driving both claim volume and severity is transactional work. Their words: the exposure has “moved from theoretical to real.”

Worth being careful here, because this one gets overstated constantly. Most professional liability policies still cover AI-related claims. Specific AI exclusions exist, but they’re narrower than the headlines suggest.

The evidence ledger

The famous number (VENDOR):LawGeex whitepaper, February 2018. 20 lawyers, 5 Enron NDAs, issue-spotting only, AI 94% against lawyers 85%. Answer key built from “the best answers of all 21 participants (including the LawGeex AI).” Never peer reviewed, never replicated, page now offline.

Where AI wins (Vals Legal AI Report, February 2026, 8 firms, vendor-adjacent):document Q&A 94.8 against 70.1, summarization 77.2 against 50.3, data extraction 75.1 against 71.1. On a 40-field credit-agreement extraction the humans lost, 61.7% against 39.2%.

Where it doesn’t (same benchmark):redlining, lawyers 79.7 against a best tool score of 65, one of only 2 tasks out of 7 where humans beat every tool. Unlabelled MFN clause: one tool confidently wrong, others found nothing. Warranty disclaimers: none found all three.

The trials (peer-reviewed):n=60, no significant quality gain, 12% to 32% faster (109 Minn. L. Rev. 147, 2024). n=137, gains that “do not appear to extend to the one transactionally oriented task” (Schwarcz et al., 2026).

Hallucination (peer-reviewed):Lexis+ AI 17%, accurate 65%. Westlaw roughly a third. Ask Practical Law 17%, accurate 19% (JELS22(2), 2025). Legal research, not contract review.

Who pays (vendor terms, verbatim):Spellbook §8.4(b)(i), Harvey §1.5, LexisNexis §12.11. Liability typically capped at 12 months of fees. Harvey: $250,000 wrong output, $500,000 data breach.

The docket:empty, 2024 to 2026. FTC v. DoNotPay, February 2025, $193,000.Moffatt v. Air Canada, 2024 BCCRT 149, CAD $812.02.

The insurers:7 of 13 carriers covering over 80% of Am Law 200 firms report rising AI claims, transactional work dominating (EPIC, May 2026). Most policies still cover AI claims.

Whoever signs it owns it

The profession already wrote the rule, and it’s the cleanest statement of the whole idea I’ve read anywhere. ABA Formal Opinion 512 says lawyers “may not abdicate their responsibilities by relying solely on a GAI tool,” and that “the lawyer is fully responsible for the work.”

The rule is about accountability, not software, and it doesn’t change based on what wrote the first draft.

Which is why I wanted to walk through this even if you never review a contract for a living, because you sign them. The vendor agreement, the lease, the terms you clicked through to use whatever you’re using right now. Your deck. Your brief on the firm’s letterhead. The contract you send a client. AI is genuinely good at finding things in those documents and telling you what they say. It’s measurably worse at telling you which part is going to hurt you. That judgment is still yours, and so is the signature.

Where I stand

None of this makes me anti-AI. I use these tools every day, and on the finding and the summarizing they save me real time. I’d just like the numbers people quote at me to survive being read.

94% sounded like a fact for eight years. It was a marketing document with a footnote nobody finished.

The short version

The 94%-against-85% comparison comes from a 2018 vendor whitepaper testing its own product on 5 NDAs with 20 lawyers, where the answer key was built partly from the AI’s own answers. Never peer reviewed, never replicated, page gone. The pressure behind it is real: legal spend at 0.43% of revenue, volumes up for 81% of departments while 55% hold flat budgets. Independent evidence says AI wins at finding, summarizing and extracting, and loses at redlining, 79.7 against 65. Two randomized trials agree the gains don’t extend to contract work. The vendors’ own terms disclaim accuracy and cap liability at around 12 months of fees. Nobody has been to court over it yet, because the risk got assigned in advance.

If you work with contracts, I’d like to hear where these tools have genuinely helped you and where you’ve caught one getting it wrong. What did the failure look like in practice? And if you know this field better than I do, tell me what I’ve missed. I’d genuinely love to know more.

If this was useful, my book Founders Who Finish is over at davesaunders.net, and my newsletter, The Build, is there too.