Skip to content

Published · 7 min read

But would a lawyer actually send that?

General Legal’s new benchmark evaluates the gap between model performance in completing discrete legal tasks and in drafting client-ready deliverables.

As any partner marking up a junior associate’s work will tell you, a legal draft can look polished, complete all the required changes, and still not be client-ready. A rubric might be able to tell us which requirements an agent met; we also need to ask whether its choices fit together and serve the client’s priorities. As frontier models continue to improve their ability to handle legal tasks, it’s becoming increasingly important to differentiate technical completion from a document that is client-ready.

Because this is a fully internal benchmark that never leaves our system, each test begins with a real General Legal client matter for which we have permission to run the tests. The agent receives the request and the information available when our attorney begins work. We evaluate its output against what GL’s attorneys actually delivered, drawing on their attorneys' judgment and expertise about the matter. The agent works in a controlled harness modeled on our production environment and access to the firm's proprietary knowledge base.

Expert legal work is a sequence of choices: which risks are absorbable, what client objectives a clause must protect (and, often, which must be traded in exchange for serving other business goals), and when the right answer belongs in the client advice rather than the redline. Because General Legal is a law firm and already has experienced lawyers making those evaluations every day, we can efficiently examine those choices in the revisions and client-facing work our attorneys actually produced. This makes domain expertise something we can test: track the individual changes an agent makes, then examine the full package for omissions, conflicting terms, or advice that doesn’t match the version an attorney released to a client.

At the same time, we can also evaluate the output holistically. Our benchmark allows for changes that reach the same outcome via other drafting conventions, but also identifies when the models add unnecessary stipulations, make strange choices that we wouldn’t release to a client or counterparty, or simply overshoot the mark, as eager models often do. Legal review is far from the checkbox exercise of fulfilling N binary requirements that some other benchmarks seem to reflect.

An added benefit of running the benchmark on real matters is that we can more accurately reflect the expansive, iterative context in which real attorneys work. Legal work unfolds inside a client relationship. Rather than arriving as a naïve, sterile task, client matters have a history: prior work, client communications and preferences, and, in negotiations, responses that can change the direction of the deal. Our benchmark draws on real General Legal matters across different practice areas, following each assignment from the initial request to the first completed work our attorney delivered to the client.

Our benchmarking agents start with the documents and client context available to the lawyer at that point. With Ansible, it can also draw on the firm’s accumulated legal knowledge and workflows in a harness modeled on production. The matter record—including follow-up requests, changes in direction, counterparty responses, and attorney revisions—helps us understand the choices behind the delivered work and ground the evaluation in a real client outcome rather than some purely academic evaluation of whether the output is technically correct.

Most models aren’t there yet

An agent is most valuable to an attorney when its draft can go out without further edits. So alongside the usual criterion pass rate, we report all-pass rate: the number of matters where an output meets every criterion. A draft that misses even one criterion still needs an attorney to find and fix the miss, and all-pass rate reflects that. We tested 13 recently released models, both open and closed source, all at High reasoning effort.

Results for 13 models at High reasoning effort. Left: share of rubric criteria met. Right: matters that met every criterion according to both LLM judges.

By criterion pass rate, models look strong. Nine of the 13 meet more than 80% of criteria, and the best, GPT-6 Astra, meets 94.7%. All-pass rate tells a different story. Nearly half the models (6 of 13) fully pass only one matter out of 10. Even the GPT-6 Astra fully passes just 6 of 10, and it's the only model to pass more than half.

The two measures also don't track each other closely. Muse Spark 1.3 and GPT-5.6 Sol meet almost the same share of criteria (83.5% and 83.1%), yet Sol fully passes twice as many matters (4 versus 2). Grok 4.7 meets 92.0% of criteria, second only to Astra, but fully passes fewer matters than GPT-5.6 Sol. GPT-6 Luna meets 85.7% and fully passes one. Getting most of a matter right is a different skill from getting all of it right, and only the second produces work a lawyer would send.

Just like in the actual practice of law, thinking harder helps

Model choice is only part of the story. We ran three GPT models with varying sizes (GPT-5.6 Terra, GPT-5.6 Sol, and GPT-6 Astra) at Low, High, and Ultra reasoning effort, and all-pass rates varied substantially within the same model family.

All-pass rate (every criterion met for both judges) by reasoning effort for three GPT models.

For the two stronger models, more reasoning paid off. GPT-6 Astra fully passes 2 matters at Low, 6 at High, and 7 at Ultra, and its jump from Low to High is the largest in our results. GPT-5.6 Sol climbs steadily, from 1 to 4 to 6. GPT-5.6 Terra, by contrast, passes just one matter at every setting.

Our hypothesis is that reasoning pays off once a model's baseline is strong enough. A legal assignment asks an agent to reconcile connected decisions: a redline may affect another clause, the client's negotiating position, and the advice accompanying the draft. Extra reasoning gives a capable model room to close those last few gaps, which is what takes a draft from nearly right to fully right. Terra's baseline may simply be too low for that.

Extra reasoning doesn't pay off equally for every model. GPT-5.6 Sol and GPT-6 Astra can both fully pass 6 matters, but at very different costs: Sol needs about 7.5M tokens per matter at its Ultra setting, while Astra needs about 1M at High. Terra, meanwhile, more than doubles its token use from High to Ultra without passing a single additional matter. More thinking helps, but some models turn it into better work far more efficiently than others.

Estimated generation cost per matter (log scale) versus all-pass rate for 13 models at High reasoning effort, based on standard API rates as of September 24, 2026.

Across models, a higher price doesn't buy better work. Costs range from about $0.07 to $6 per matter, and the most expensive models aren't the strongest. Fable 5.1 costs about $6 per matter and fully passes 2, and Claude Opus 5 costs about $4 and passes 1. The best result isn't the cheapest either. GPT-6 Astra fully passes 6 matters at about $3.4 each.

New releases are changing the picture as well. Claude Opus 5.5 costs less than Opus 5 and fully passes three times as many matters. GPT-6 Sol costs less than half as much as GPT-5.6 Sol, but passes one fewer matter. The pattern isn't clean yet, but newer models are generally getting cheaper, and some are getting better at the same time.

Looking ahead

Despite its newly public status, the General Legal Benchmark will retain the same simple target as when we developed it for internal use: if the assignment arrived from a client today, is this what the lawyer would actually send?

In the near future, we’ll share our complete benchmark on how close the models and harnessing come to producing work that not only meets technical requirements but also generates client-ready work.

We will also share fuller comparisons of performance, inference costs, drafting speed, and uplift from domain-specific harnessing. Stay tuned for where models excel at freeing attorneys for higher-level work, where differentiated knowledge sets add particular value, and which decisions still call for an attorney.

Key takeaways
  • General Legal's new benchmark evaluates whether AI-drafted legal work is truly client-ready, not just technically complete.
  • The benchmark tests agents on real client matters with actual attorney work as the gold standard, including access to the firm's proprietary knowledge base.
  • Legal work requires judgment about trade-offs, priorities, and strategic choices that extend beyond simply fulfilling a checklist of requirements.
  • The benchmark tracks both individual changes and holistic quality, identifying when models add unnecessary terms or make choices attorneys wouldn't release to clients.
  • All-pass rate measures the percentage of matters where output meets every criterion, reflecting that even one miss requires attorney intervention.
TL;DR
The gap benchmarks missTechnical completion differs from client-ready work; a draft can meet all requirements and still need substantial revision before release.
How the benchmark worksAgents receive real General Legal client matters and are evaluated against what GL attorneys actually delivered, with access to the firm's knowledge base.
What it evaluatesBoth individual changes and the complete package, checking for omissions, conflicting terms, unnecessary additions, and choices that don't serve client priorities.
Real-world contextTests reflect iterative legal work within ongoing client relationships, including prior work, preferences, and negotiation dynamics rather than sterile isolated tasks.
All-pass rate metricMeasures matters where output meets every criterion without attorney edits, since even one miss requires human intervention to find and fix.
Current model performanceTesting of 13 recent models at high reasoning effort shows most aren't yet producing consistently client-ready work.
FAQs

What makes this benchmark different from other legal AI evaluations?

It evaluates complete client-ready work rather than isolated task completion, using real General Legal matters and comparing agent output against what experienced attorneys actually delivered to clients. The benchmark examines both technical accuracy and holistic judgment calls that make legal work fit for release.

Why does all-pass rate matter more than regular pass rate?

An agent draft is only truly valuable if it requires no further attorney edits. If even one criterion is missed, an attorney must review the entire document to find and fix the problem, eliminating much of the efficiency gain. All-pass rate reflects this real-world constraint.

How does the benchmark handle the subjective nature of legal judgment?

By testing on real client matters where General Legal attorneys have already made and documented their judgment calls, the benchmark grounds evaluation in actual professional decisions rather than abstract standards. It allows for different drafting approaches that reach the same outcome while flagging choices that experienced attorneys wouldn't make.

What context do the AI agents receive when being tested?

Agents start with the same documents and client context available to the attorney at the beginning of the matter, including prior work, client communications and preferences, and access to General Legal's proprietary knowledge base through a production-like harness.

Are current AI models passing this benchmark?

No. Testing of 13 recently released models at high reasoning effort shows that most aren't yet consistently producing client-ready work that meets all criteria without attorney revision.