Harvey LAB-AA v1.1 Adds Hallucination Checks to Its Legal-Agent Evaluation
Harvey LAB AA v1.1 Adds Hallucination Checks to Its Legal Agent Evaluation October 8, 2026 Harvey LAB AA v1.1 updates the scoring methodology for the Legal Agent Benchmark (LAB)...
By AI Engineering Team
Harvey LAB-AA v1.1 Adds Hallucination Checks to Its Legal-Agent Evaluation
October 8, 2026
Harvey LAB-AA v1.1 updates the scoring methodology for the Legal Agent Benchmark (LAB), an evaluation of AI agents performing real-world legal work. The revised benchmark checks every deliverable for hallucinations against its source documents, uses a three-judge panel for rubric grading, and introduces the Hallucination-Gated All-Pass Rate.
A task receives credit for this headline metric only when its deliverables satisfy every rubric criterion and contain no material hallucinations.
The evaluation covers 120 private legal tasks created by Harvey. The tasks span corporate mergers and acquisitions, capital markets, tax, litigation, bankruptcy, and other legal practice areas. Future updates are intended to account for additional factors lawyers value, including usability characteristics such as style and tone.
Headline results
Grok 4.7 (xhigh) leads Harvey LAB-AA v1.1 with a 9.4% Hallucination-Gated All-Pass Rate. Muse Spark 1.3 (max) follows at 8.9%, and GPT-6 Astra (max) records 8.6%. The remaining listed results are:
- GPT-6.1 Sol (max): 6.9%
- Claude Fable 5.1 (max, with fallback): 6.4%
- Kimi K3 (max): 5.3%
- Claude Opus 5.5 (max, with fallback): 4.2%
The Claude models were run with Anthropic's fallback enabled, but the fallback was not used.
Accounting for hallucinations substantially changes the rankings. Without the hallucination gate, Muse Spark 1.3 would lead with a 26.7% All-Pass Rate. However, two-thirds of those passes contain a material hallucination. GPT-6 Astra retains nearly all of its passes, declining from 8.9% to 8.6%, and moves from joint 10th place on All-Pass Rate to third place on the hallucination-gated measure. GPT-6.1 Sol declines from 7.5% to 6.9%.
Across the tested models, more than 60% of otherwise passing results contain a material hallucination. Several models score 0% after hallucinations are taken into account.
Most models complete a large majority of rubric criteria. Sixteen models pass between 85.6% and 96.0% of criteria. The rubrics assess each deliverable as a whole rather than focusing only on criteria selected for difficulty. Criterion Pass Rate is therefore reported separately from hallucinations to show rubric coverage and source grounding independently.
Changes in Harvey LAB-AA v1.1
Version 1.1 is not directly comparable with previously published v1.0 results because the grading and scoring methodology changed.
Hallucination auditing
Every submitted deliverable is audited against the task's source documents through a two-stage review. A task with no usable submission receives a score of zero and is not audited.
During the first stage, a judge flags possible hallucinations in three categories:
- Contradictions of the source materials
- Fabricated source content
- Specific assertions that lack support in the sources
The same judge then rechecks every flag against the task's source documents and classifies upheld flags as material or minor. Flags that do not hold up are dismissed.
Hallucination-Gated All-Pass Rate
The headline metric is now the Hallucination-Gated All-Pass Rate. A task counts only when all deliverables pass every rubric criterion and contain no material hallucination.
Criterion Pass Rate and material hallucinations
The benchmark reports the share of rubric criteria passed before the hallucination gate. This figure is averaged across the three judges and pooled across all criteria. It is reported alongside the mean number of material hallucinations per checked task.
Three-judge rubric panel
Rubric criteria are graded by three LLM judges:
- GPT-6 Sol
- Grok 4.7
- Claude Opus 5.5
Their verdicts are averaged, replacing the single judge used in v1.0. The judges were reviewed for self-preference, and the evaluation found minimal preference for their own model families. Averaging the three results is intended to reduce the effect of any individual bias.
Dataset update
Version 1.1 uses Harvey's latest private dataset, v1.1.0, which includes improvements to the tasks and criteria.
How Harvey LAB-AA differs from Harvey's LAB
Harvey LAB-AA is an independent reimplementation of Harvey's evaluation. It differs from the original in several ways:
- Models run on the Stirrup agent harness, which supports features such as context compaction instead of failing when a context limit is reached. The implementation also uses simplified agent and judge prompts authored for the evaluation.
- Harvey's custom tools and document-generation skill scripts, including
pptxanddocxtools, are not included. Instead, the evaluation provides a simple code-execution tool to measure raw model ability. - Deliverables must use the exact filenames specified by the task. Fuzzy matching is not used when a model produces an incorrect filename.
Hallucination methodology
Human preference studies conducted by Harvey found that experts frequently identified hallucinations as a primary factor when choosing between two otherwise comprehensive answers. The LAB-AA check focuses on errors that could materially affect legal work and uses a conservative flagging approach.
The check distinguishes unsupported task-specific claims from general legal knowledge. General legal knowledge, including case law or statutes beyond the source files, is outside the scope of the check and does not by itself count as a hallucination.
A hallucination is a claim in a work product that is not supported by the task's source documents. Each hallucination falls into one of three categories:
- A contradiction of the source materials
- Fabricated source content
- A specific assertion with no support in the record
Each upheld hallucination is classified as material or minor. A material hallucination could mislead a reader on a substantive issue, such as an incorrect contractually required date. A minor hallucination is a genuine error that is unlikely to meaningfully affect the legal interpretation of a deliverable.
Only material hallucinations affect the headline score. One or more material hallucinations reduces a task's Hallucination-Gated All-Pass Rate contribution to zero. Criterion Pass Rate is measured before the hallucination gate. Material and minor hallucination counts are reported per task on which the check ran, and minor hallucinations do not affect scores.
Three-step hallucination check
- Find possible hallucinations: The judge compares the submission with the task sources and flags sections that may contain unsupported claims.
- Confirm or dismiss: The same judge checks each flagged section against all input source documents, either upholding it with a severity or dismissing it.
- Gate the score: Any confirmed material hallucination reduces the task's gated score to zero.
Illustrative tax-compliance example
One example task involved a tax-compliance review. The source materials stated that a $6 million principal repayment reduced the same loan from $220 million to $214 million during 2023.
A submission listed the debt as “$220M term loan + $214M notes”. The judge confirmed this as a material hallucination because the two figures represented the opening and closing balances of one loan, not two separate instruments.
The same submission included the calculation “$6,400,000 * 293/365 $5,136,986”. The source calculation was approximately $5,138,000, and the stated arithmetic equals approximately $5,137,534. The judge upheld this as a minor hallucination.
A separate submission stated, “The $960,000 overpayment was credited to 2024”. The source requested confirmation before filing but also expressly showed $960,000 applied to 2024 estimated tax. The judge dismissed the flag as supported by the source.
In this example, the material hallucination makes the task's Hallucination-Gated All-Pass Rate 0%. Criterion Pass Rate remains based on rubric judgments. In an illustrative three-judge result, two of three judges passed every criterion, producing a 67% All-Pass Rate, while a material hallucination reduced the gated rate to 0%. Of 30 individual judge verdicts, 27 passed, producing a 90% Criterion Pass Rate.
These examples come from the same tax task. The material and minor examples share one submission, while the dismissed example comes from another submission. They are illustrative excerpts, not a complete task audit.
Hallucinations by model
The GPT-6 models record the fewest material hallucinations in the full evaluation. GPT-6 Astra averages 0.03 material hallucinations per task, with four hallucinations across all 120 tasks. GPT-6 Sol averages 0.07 per task, with eight hallucinations across the full set.
Gemini 3.8 Flash records the highest average in the launch set, at 13.96 material hallucinations per task. Muse Spark 1.3 passes the highest share of criteria, 96.0%, but averages 1.68 material hallucinations per task, compared with 0.03 for GPT-6 Astra. Every open-weights model averages at least 2.09 material hallucinations per task.
Criterion completion and source grounding are different capabilities. The results also indicate that hallucination frequency depends more on the model than on the legal practice area.
Selecting the hallucination checker
GPT-6 Sol (high) is used for both stages of the hallucination check. It is separate from the three-judge rubric panel.
Six candidate hallucination judges were compared at high reasoning effort:
- GPT-6 Sol
- GPT-6 Luna
- Grok 4.7
- Claude Opus 5.5
- Claude Sonnet 5.5
- Gemini 3.8 Flash
The comparison used the same 20 tasks and deliverables from eight evaluated models. GPT-6 Sol and GPT-6 Luna generally identified more material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash identified substantially fewer. Claude Opus 5.5 fell between Grok 4.7 and Claude Sonnet 5.5 in every model row.
Across the comparison, Opus upheld 99 material hallucinations, Grok upheld 219, Sonnet upheld 57, and GPT-6 Sol upheld 470. GPT-6 Sol identified more material hallucinations despite the conservative approach, which excludes general legal knowledge not contained in the source documents.
All six checkers found no material hallucinations in GPT-6 Astra's outputs. GPT-6 Sol ranged from 0 to 0.20 material hallucinations per task for GPT-6 Sol's outputs. Both models were among those with the fewest material hallucinations under every checker.
The fixed-subset comparison does not change the full 120-task leaderboard, which uses GPT-6 Sol (high) as the hallucination checker.
Hallucination-Gated Near-Pass Rate
The Near-Pass Rate allows tasks that miss one or two criteria, while still assigning a zero to any task with at least one material hallucination.
GPT-6 Astra rises from an 8.6% Hallucination-Gated All-Pass Rate to 20.3% when one criterion may be missed and 31.7% when two may be missed. GPT-6.1 Sol rises from 6.9% to 20.3% and 29.6%, respectively. Grok 4.7 records 15.3% and 23.1% in those two bands.
The next largest gains are recorded by GPT-6 Sol, which rises from 3.6% to 18.1%, and Claude Sonnet 5.5, which rises from 2.8% to 16.9%.
Models whose apparent passes are removed by hallucinations gain little from allowing missed criteria. GLM-5.3 reaches only 2.2% with two missed criteria, while Gemini 3.8 Flash remains at 0%.
Cost
The highest-scoring models are not the most expensive. Grok 4.7 leads at approximately $9.50 per task, less than half the cost of Claude Fable 5.1 at approximately $21.70 per task. Muse Spark 1.3 ranks second for approximately $4.20 per task.
Evaluation cost is divided into input, cache-hit, cache-write, reasoning, and answer-token costs where canonical token counts are available.
Token usage
Generating more output tokens does not necessarily produce a higher score. GPT-6 Astra scores 8.6% while using approximately 81,000 output tokens per task, less than half of Grok 4.7's approximately 180,000 tokens.
The three Claude models generate the most output, approximately 202,000 to 562,000 tokens per task, while scoring between 2.8% and 6.4%.
Speed
Stronger models generally take longer to run. Estimated decoding time is approximately 33 minutes per task for both Grok 4.7 and Claude Fable 5.1. Muse Spark 1.3 records the second-highest score at approximately 12 minutes per task.
These estimates exclude time to first token and other overhead.
Turns
The highest-scoring models do not necessarily use the longest agent loops. Grok 4.7 averages approximately 63 turns per task, and GPT-6 Astra averages approximately 54. Claude Sonnet 5.5 uses the most turns, approximately 179 per task, while scoring 2.8%.
The turn count is a rough proxy for the number of actions, tool calls, and iteration cycles an agent uses to complete a task.
Open-weights model compute proxy
For open-weights models, the compute proxy is calculated as:
Active Parameters (billions) × (Input Tokens / 5 + Output Tokens) / 1,000,000
Input tokens are downweighted to reflect the lower compute cost of prefill compared with generation. Lower values indicate greater compute efficiency. In the chart, models in the upper-left region combine higher scores with lower estimated compute use.
Example tasks and submissions
The public task set includes representative tasks covering several legal workflows:
- M&A change-of-control analysis
- Deposition outline
- Arbitration agreement redline
- M&A disclosure schedules
- Commercial lease review
One M&A task asks the model to review acquisition data-room contracts and an internal memo for change-of-control and assignment provisions, then prepare a comprehensive deal-team report. The required output is coc-analysis-report.docx.
The reference files include:
apex-distribution-agreement.docxapex-msa.docxdeal-overview-memo.docxdeal-summary-memo.docxfirst-continental-credit-agreement.docxgreat-lakes-credit-agreement.docxmeridian-dpa.docxmeridian-license-agreement.docxnovabridge-partnership.docxnovabridge-supply-agreement.docxorion-subscription-renewal.docxpinnacle-ecommerce-agreement.docxpinnacle-license.docxridgeline-10k-excerpt.docxsolara-deferred-comp-plan.docxterranode-isa.docxterraverde-lease-agreement.docxwebb-employment-agreement.docxwellstone-manufacturing-agreement.docx
The example set also contains model-produced deliverables for Claude Opus 5.5 (max with fallback), GPT-6 Astra (max), GPT-6.1 Sol (max), Grok 4.7 (xhigh), and Muse Spark 1.3 (max).
Resources
- The Harvey LAB-AA evaluation page contains the leaderboard and full results.
- The methodology documentation describes the implementation, including agent and rubric-grading prompts.
- Harvey's original LAB announcement describes the benchmark and its design.
- A public set of representative tasks is available on GitHub.
- Harvey LAB-AA runs on the Stirrup open-source agent framework.
Results in this article reflect data available on October 8, 2026.