AI News HubLIVE
In-site rewrite6 min read

Show HN: LitigationBench. A Litigation Task-Based AI Benchmark

LitigationBench is a benchmark from Litco for evaluating language models on litigation tasks. Each model runs tasks twice: without and with Litco's safeguards, with both scores and failures published. Special task sets test practitioner indistinguishability, cert-QP framing, AI-isms, case characterization, calendaring, and candor. Models that fabricate case law lose routing eligibility and incur score penalties. The methodology is transparent, with private task sets to prevent overfitting.

SourceHacker News AIAuthor: ybathaee

LitigationBench

by Litco

A benchmark for language models on litigation work. Every model runs the same tasks twice: once without Litco’s safeguards, once with them active inside Litco’s production agent. Litco publishes both numbers, the gap between them, and every failure, including failures by the models inside the product.

loading…

Leaderboard

Click a column to sort (arrows mark the direction that counts as better); click a row for both settings side by side and the row’s serving pins. Models missing a setting in this preview sit below the scored rows with the reason stated.

The composite is the run’s quality score for the setting shown. Litco publishes the score before penalties beside the composite, so readers can see the effect of those penalties for themselves. The fabricated- authorities and false-premises columns are shaded per cell: green for zero, amber for an adopted false premise, red for a fabrication; the full flag detail is in each row’s expanded panel and in the candor matrix below. Cost is the metered provider bill for the tasks shown. The self-hosted row runs on Litco’s own hardware; its cost column is priced at OpenRouter’s current rates for the same open-weight model, qwen/qwen3.6-35b-a3b, so every row’s cost rests on the same basis.

  • Without safeguards, the cost column shows the bill for all 59 tasks rather than a fraction of a cent per task.

Fabricated authorities, without Litco’s safeguards and with them.

Every fabricating model reports zero fabricated authorities on its ordinary work with Litco’s safeguards active. The handful that remain surfaced only when the run deliberately pressed a model with a false premise or an invented authority against the live corpus. The chart pairs each model’s count without Litco’s safeguards against the same model’s count with them active. A green tick marks a measured zero, and an amber bar counts fabrications that happened on those provoking tasks. Lower is better, and zero is the goal.

without Litco’s safeguards with Litco’s safeguards, measured zero with Litco’s safeguards, on the tasks written to provoke fabrication

hover a bar for the exact count · sorted by fabrications without safeguards

Each model’s cost and score.

Each dot is one model. Farther right costs more per task, and higher up scores better on the tasks. The self-hosted model runs on Litco’s own hardware and is priced at OpenRouter’s rates for the same model, so every dot rests on the same cost basis. Click a dot to compare two models side by side.

Vertical axis: Horizontal axis:

hover for detail · click to pin into the comparison · models that invented authority carry a red ring and lose Litco routing eligibility

Which models invent case law, and how often.

Counts of fabricated authorities per model over the 59-task battery without Litco’s safeguards. These are inventions of legal authority by the model running alone.

Failures under pressure, model by model.

The candor tasks press each model with invented premises, questions that lack a real legal response, and pressure to produce authority for a position no authority supports. The chart counts three failures for each model: fabricated authorities, adopted false premises, and hedges under pressure. Red bars show the model without safeguards, and blue bars show the same model inside Litco.

Per-category skill scores, inside Litco

Six skills, graded per task and rolled up per model: candor under pressure, faithful case reading, treatment awareness, procedural competence, appellate issue framing, and reasoning and drafting.

score colors, shared by every tinted chart on this page: 85% and up 60% to 84% below 60%

Special task set · practitioner indistinguishability

Platform introductions against filed ones, blind.

This task set takes the introduction section of a real OT2025 Supreme Court merits brief and has each model rewrite it from the same record. A blind three-judge panel, cross-vendor and in both presentation orders, first tries to spot which one a machine wrote, then, with no AI framing at all, simply says which introduction it prefers. Detection sits close to a coin flip for most models. On preference, several models’ rewrites beat the filed brief most of the time.

above the 50% coin flip: the panel preferred the model’s introduction more often than the filed one at or below the coin flip: the panel liked the filed introduction at least as often

Special task set · cert-QP framing

Ranking question-presented drafts, pairwise.

This task set hands each model a granted certiorari petition and has it draft the question presented. One petition was granted after every model’s training cutoff and counts double. A judge then compares every pair of drafts head to head rather than scoring each draft alone, and the bars show how often each model’s draft won its comparisons.

Special task set · AI-isms

How machine-written each draft reads.

The benchmark scores each model’s drafting output against a public list of recognizable AI writing tells: tic words, spaced em-dashes, stacked verbless fragments, the “It’s not X. It’s Y.” reveal, and the rest of the list, reported as tells per 1,000 words rather than pass or fail. Every model wrote the same three pieces: a brief argument, a client letter, and a research memo, each with the same length target and all of the law supplied in the prompt, so every rate rests on a comparable amount of text. These rates now count against the leaderboard: a model’s drafting score is reduced by its rate of AI writing tells, up to fifteen points. Lower is better, and zero is the goal.

Some models wrote too little in this task set to measure fairly. Their rows show “not enough text to score” instead of a number.

What gives each model away

One card per model, heaviest accent first. Each chip is one tell counted in that model’s three drafts, with how often it fired; the chips that name a construction, such as the “It’s not X. It’s Y.” reveal, quote the pattern being counted, not the model’s sentence.

Special task set · case characterization

Reading an opinion: holding, disposition, facts.

This task set gives each model an opinion and grades three things: separating the holding from dicta, stating the disposition correctly, and getting the facts right. Each bar is the share of those calls the model got right, graded against a reading of each opinion that a person checked by hand. Each dot places one model inside Litco, best to worst, on an axis zoomed to where the scores actually sit.

Special task set · calendaring, FRCP/local rules

Deadlines computed under the Federal Rules.

Calendaring tasks hand each model a trigger date and a filing deadline to compute under the Federal Rules. Every deadline in the battery is checked against a computer program that applies Rule 6(a)’s counting rules directly, so the grading never depends on another model’s arithmetic. Each dot places one model inside Litco, best to worst, on an axis zoomed to where the scores actually sit.

Special task set · candor, across three settings

The candor matrix, three settings.

The candor tasks press each model with invented premises, questions the record cannot resolve, and pressure to cite authority, across three settings: without Litco at all, with Litco active but no real corpus mounted, and with Litco active against the live 3.6-million-opinion litlex corpus under deliberate pressure. Each cell below counts fabricated authorities and adopted false premises for that model in that setting. Hedge counts ride along for information only and never turn a cell red.

no failures: neither a fabrication nor an adopted premise fabricated authority (F) adopted false premise (P) H = hedge, informational only

Head to head

Any two models, axis by axis, in the mode selected on the leaderboard. Pick from the dropdowns or click points on the scatter.

VS

Method

What a run is

A model is dropped into the production Litco matter agent (the same agent loop, tools, and verification stack Litco’s customers run) and given the task battery against a seeded litigation matter and the production case-law corpus. Nothing about the harness is model-specific: every model sees identical prompts, tools, and iteration budgets. Spend is metered per call.

The two modes

Battery v2 measures every model twice. With Litco safeguards is the production configuration: the agent with its verification stack active, over a 42-task battery. Without safeguards is the model on its own over a 59-task battery that includes the fabrication probes: the stack observes and records but never intervenes. The per-model difference between the two composites is printed on the chart above. The two batteries share their core tasks; composites are compared across modes as the run reports them, and the battery sizes are stated wherever the numbers appear.

Candor penalties and routing eligibility

A model that invents case law, citing an authority that does not exist, wears a red flag on every surface of this page and loses twenty points from its composite for each fabrication. A model that adopts a false premise embedded in a question loses ten points for each adoption. Litco floors every composite at twenty, however many penalties a model accumulates in a run. A model that fabricates also loses Litco routing eligibility for that mode, whatever its score, because an invented citation is the one failure a legal tool may never let through. A model’s drafting score is also reduced by its rate of AI writing tells, measured in this page’s AI-isms task set, up to fifteen points. Litco publishes the score before penalties beside the composite, so readers can see the effect of each penalty for themselves.

IF THERE IS ANY FIXED STAR IN OUR CONSTITUTIONAL CONSTELLATION IT IS THAT NO OFFICIAL HIGH OR PETTY CAN PRESCRIBE WHAT SHALL BE ORTHODOX IN POLITICS NATIONALISM RELIGION OR OTHER MATTERS OF OPINION

IF THERE IS ANY FIXED STAR IN OUR CONSTITUTIONAL CONSTELLATION IT IS THAT NO OFFICIAL HIGH OR PETTY CAN PRESCRIBE WHAT SHALL BE ORTHODOX IN POLITICS NATIONALISM RELIGION OR OTHER MATTERS OF OPINION

Infrastructure failures

A task whose model calls moved zero tokens is an infrastructure failure: it lands in an infra count, never in a quality axis. A row where most tasks failed that way publishes as quarantined with the failure stated, rather than as a low score. One assisted row is quarantined in the current preview, and the leaderboard says so on the row.

Serving policy and pinning

Public rows run through OpenRouter pinned to a named provider at a disclosed precision, with zero-data-retention routing requested at call time; two models were not ZDR-routable through OpenRouter and ran on the vendor’s direct API, which their rows disclose. Self-hosted rows run on Litco-operated hardware and are labeled as such. Because scored traffic goes through public endpoints rather than Litco infrastructure, a model vendor can independently verify the serving conditions of its own row. Lane or precision changes count as new rows, never silent edits.

Hold-out policy

Litco publishes the methodology and keeps the task set private, because published tasks invite tuning to them, and a tuned-to benchmark stops predicting real work. The battery is versioned; when a battery version is retired, its tasks are disclosed and a fresh version replaces them. Scores are only ever compared within a battery version.

Preview status

This page currently renders the run’s own reported numbers from its live progress feed, marked PREVIEW. Judging is still in progress for the per-category skill breakdowns, the question-presented pairwise comparisons, and the brief-introduction indistinguishability test; latency percentiles and the final uniform rescore land with the final ingest, which replaces this data, and the changelog will say so.

What the assisted mode is measuring

The verification stack the assisted rows run under is the shippe

[truncated for AI cost control]