Skip to content

AI Applications Leaderboard

AI applications tested on real-world legal work.

Results updated August 2026

  • Substance: Whether the legal work is correct, complete, and responsive. Measured as task pass rate, the percentage of tasks for which every applicable criterion was satisfied.
  • Form (1 to 3): whether a lawyer would want to use the delivered work as written, including clarity, structure, and right-sizing.

Contract workflow leaderboard

Contract Workflows results: task pass rate and form for each application.
Claude Cowork (Fable 5)Claude33.9%2.61
Claude Web App (Fable 5)Claude32.3%2.53
Claude Web App (Sonnet 5)Claude29.0%2.57
ChatGPT Web App (GPT 5.6 Sol)ChatGPT27.4%2.75
ChatGPT CodexChatGPT25.8%2.77
Gemini Web App (Gemini 3.1 Pro)Gemini14.5%2.33
Microsoft Copilot Web AppCopilot12.9%2.39

Key takeaways

  • Claude led on substance. Claude Cowork recorded the highest contract-workflow pass rate, while the Claude Web App with Fable 5 led data extraction and the combined benchmark.
  • ChatGPT Codex led on form in both categories. It produced the clearest, best-structured and most appropriately sized deliverables.
  • Gemini 3.1 Pro's substantive pass rate trailed every tested Claude and OpenAI application. That pattern held in both contract workflows and data extraction.
  • Microsoft Copilot recorded the lowest substantive pass rate in both categories. It also had the lowest form score overall across the two categories.

Methodology

What we benchmark

The Applications leaderboard evaluates complete AI products through their normal user interfaces, as a customer would use them. It covers legal AI applications, general-purpose AI applications, and AI-native law firms combining human and AI capabilities.

We evaluate the application as a product, not only the underlying model. This includes its legal analysis, document handling, configuration, interface, and ability to produce work that is usable as delivered.

Every application is evaluated by Legal Benchmarks under our control. Results are never self-reported.

The task set

Applications complete demanding, real-world legal assignments contributed by practitioners in the Legal Benchmarks community. The set covers two principal categories:

  • Contract workflows: producing or amending contract language in response to a lawyer’s instructions, including basic clause drafting, template-based drafting, and bespoke drafting.
  • Data extraction: locating and reporting information from individual documents or sets of related documents, including imperfect source materials such as scans, images, and files containing tracked changes.

Tasks include native Word files, PDFs, scans, images, and documents with tracked changes. These files are uploaded in their original formats, so the application’s ability to ingest and work with them forms part of the evaluation.

Practitioner-authored tasks make up the core of the set. Synthetic tasks supplement them by targeting failure modes identified in earlier Legal Benchmarks research and published legal-AI benchmarking work. Every synthetic task is validated by a human expert before entering the benchmark.

Each task has a fixed set of binary pass/fail criteria authored and reviewed by lawyers. The criteria specify what a satisfactory answer must include and avoid while accommodating different approaches that would be professionally defensible.

How applications are evaluated

Applications are tested through their ordinary product interfaces and configured as a typical customer would use them, unless a different configuration is expressly disclosed.

The complete task set is run twice. Each run is conducted and collected independently, and the application’s score is calculated across the two runs. This tests whether the product can produce reliable work more than once rather than rewarding a single successful output.

Judges assess the application’s final submission, including its written answer and any files delivered. They do not grade the reasoning or process that produced it.

What we measure

Applications are assessed across three dimensions: substance, form, and product experience.

Substance

Substance measures whether the legal work is correct, complete, responsive, and appropriately candid about uncertainty.

We report:

  • Criteria pass rate: the percentage of applicable lawyer-authored substance criteria satisfied.
  • Task pass rate: Task pass rate is the percentage of tasks for which every applicable substance criterion was satisfied, averaged across the two independent runs.

A task passes only when all applicable substance criteria pass. Strong performance on some criteria does not offset a material omission or error elsewhere in the task.

Form

Form measures whether the output is ready to use as delivered. Every output is rated from 1 to 3 on task-agnostic rubrics covering:

  • Clarity
  • Length
  • Structure and formatting

A score of 1 means the output needs rework before use, 2 means it is usable with edits, and 3 means it can be used as delivered.

Substance and form are kept separate. A polished answer may still be legally unreliable, while a substantively correct answer may require substantial editing before use.

Product experience

We also assess the practical experience of completing legal work through the application, including usability and document handling. Application performance depends on the complete product experience, not only the quality of its final text.

Grading and human validation

Substance is graded against each task’s fixed binary criteria by cross-family LLM judges. Judges apply the criteria as written and are not asked to reward a preferred drafting style.

Form is assessed by a cross-family panel applying the same task-agnostic rubrics. Panel ratings are combined into a consensus score.

The grading process includes the following safeguards:

  • Disagreement tracking: Judge disagreements are recorded for review.
  • Human adjudication: Any substance disagreement that could change whether a task passes is escalated to a qualified lawyer for a final, binding determination. A maximum-spread disagreement on form is escalated in the same way.
  • Blind review: The lawyer does not know which application produced the output or which judge issued each verdict.
  • Human calibration: Practitioners spot-check judge decisions on a sample from every run. Confirmed corrections are incorporated into the published scores.
  • Favoritism monitoring: We monitor whether graders systematically favour applications using models from the same provider and publish form sub-scores so readers can inspect relevant patterns.

Human determinations take precedence over automated verdicts and are preserved for audit.

Results and rankings

Results are reported separately for Contract Workflows and Data Extraction, alongside overall figures. Substance and form remain separate, and clarity, length, and structure and formatting are reported individually because applications can fail in different ways.

Rankings represent a point-in-time evaluation and are refreshed when the benchmark is rerun.

Certification and disclosure

Certifications are awarded quarterly based on performance. During an active quarter, participating legal AI applications are evaluated as contenders, with results confirmed when the quarter closes.

Named results for general-purpose applications and open-source references may be published on the leaderboard. A legal AI application’s name and result become public only where the vendor opts in or accepts a publicly awarded certification.

Public information may include:

  • Named results for general-purpose applications and open-source references
  • The aggregate Legal AI average
  • Published certifications and badges
  • The evaluation period and other context required to interpret the results
  • Anonymised output excerpts that cannot be traced to the application that produced them

The following remains confidential for participating legal AI applications:

  • The vendor’s private report
  • Per-task results and grading records
  • Failure-mode diagnostics
  • Improvement analysis
  • Individual results belonging to other legal AI applications

Private reports do not identify competitors. Comparative information about other legal AI applications is presented only through the aggregate Legal AI average.

Benchmark security

The task instructions, reference documents, answer keys, and per-task criteria are not publicly distributed. Releasing them would allow applications to be trained or tuned against the benchmark and would weaken its value as an independent evaluation.

Vendors can submit an application through Get benchmarked, after which Legal Benchmarks conducts the evaluation on their behalf.

Limitations

  • Coverage: The task set is English only and currently leans towards commercial, IP, employment, M&A, and competition work, with a US and UK concentration.
  • Single assignments: Each task consists of one instruction and one submission. The benchmark does not yet test iterative refinement, multi-turn workflows, or longer matters.
  • Point-in-time results: Applications change frequently. A result from one evaluation period may not represent later product versions.
  • LLM judging: Automated grading does not establish perfect agreement with lawyers across the full leaderboard. We mitigate this through disagreement escalation, human adjudication, and practitioner spot-checking.
  • Limited consistency measurement: Each task is completed twice. This provides some evidence of repeatability but does not measure the full range of output variance users may encounter.

Don’t see an application?

Tell us which legal AI applications you want benchmarked. You can submit more than one.

Get benchmarked for the next quarterly awards