LegalOn benchmark finds general-purpose AI models still fail on common contract review tasks

LegalOn's 2026 benchmark tested 11 AI models on 3,282 contract reviews and found general-purpose models often miss precise legal standards. The top-scoring system, LegalOn, outperformed the best GPT model by over 400 ELO points and completed reviews in 2.3 seconds versus 40.4 seconds.

Categorized in: AI News Legal
Published on: Jun 22, 2026
LegalOn benchmark finds general-purpose AI models still fail on common contract review tasks

General-purpose AI models continue to fail on important contract review tasks, missing precise legal standards even when they identify the right topic. LegalOn tested 11 AI models across 3,282 head-to-head reviews and 21 precision-critical guidelines, releasing the results in its 2026 Contract Review Benchmark. The evaluation shows that the software system wrapped around a model-what LegalOn calls the harness-can matter more than the model name itself.

"The model cannot just sound fluent. It has to determine whether the contract meets the standard," the benchmark report states. "In many cases, the relevant failure is not what the contract says, but what it doesn't say."

Where general-purpose models stumble

The provisions tested were common contract review issues: assignment rights, PHI ownership language, NDA purpose clauses, SOW incorporation requirements, and manuscript review timelines. Wrong answers on these can create real legal or business risk. The benchmark found that general-purpose models often identified the right topic but missed the legal standard.

Finding an assignment clause is not enough if the guideline requires an unconditional assignment right with no consent requirement. PHI handling obligations are not the same as an express PHI ownership acknowledgment. A manuscript review right does not satisfy a guideline if the review period is too short. A provision that satisfies one part of a two-part requirement does not satisfy both. These failures stem from what the contract omits, not just what it includes.

The harness changes the result

A general-purpose model reviewing a full contract in one broad pass performs a different task than a system that breaks the review into structured, provision-level checks. LegalOn's approach ties each check to a specific guideline and a specific part of the contract. Contract review is many small tasks running together: Is the clause present? Is the required statement included? Is the number within the acceptable range? Are both conditions met? Does the SOW incorporate the MSA?

LegalOn ranked first across all 21 provision types in the benchmark. Its ELO score was 87 points above the next closest model and more than 400 points above the best GPT model tested. The confidence interval did not overlap with any tested model. Speed also diverged sharply: LegalOn completed a full review in 2.3 seconds, while Claude Opus 4.6, the strongest general-purpose model on speed, averaged 40.4 seconds per contract.

How the benchmark was built

For each contract and provision, two reviews ran side by side: one from LegalOn and one from a general-purpose AI model. The baseline models received the full contract and all guidelines at once, returning MET or UNMET determinations in a single pass. An independent LLM judge, separate from the models tested and blind to authorship, assessed which review was more accurate, complete, and useful. Judging criteria included correctness, evidence quality, article identification, completeness, and reasoning quality.

To control for position bias, every comparison was run twice with the order reversed. A result counted as a win only when the same system was preferred in both orderings. If the preference flipped, the result was treated as a tie. Legal experts also validated a sample of the judge's outputs against professional legal standards.

For professionals building skills in this area, resources like the AI Learning Path for Paralegals cover the kind of structured document review and contract analysis that benchmarks like this measure. AI for Legal Professionals Courses offer additional training on evaluating and using AI tools in legal workflows.

Why this matters for legal professionals

Foundation models will keep improving, but legal AI should be evaluated on legal tasks, not on general model reputation. The relevant question for a legal team is not which model sits underneath a product. It is whether the product can reliably apply a legal standard to the details of a contract. No benchmark replaces testing a system on your own contracts, playbooks, and risk standards. But a serious evaluation can show where general-purpose models perform well, where they fail, and what product architecture is needed to turn model capability into dependable legal work.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)