Skip to content
AI News HubLIVE
Original source2 min read

Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus

Summary

A new arXiv paper shows that LLM judge consensus is less reliable than it appears because judges make correlated errors. In a ten-judge panel, the average pairwise error correlation was 0.21, making the panel roughly equivalent to 3.5 independent judges; ignoring shared errors changed significance in up to 28% of comparisons.

SourcearXiv AIAuthor: Elias Hossain, Niloofar Yousefi, Ser-Nam Lim
Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

[Submitted on 18 Sep 2026]

Title:Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus

View a PDF of the paper titled Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus, by Elias Hossain and 2 other authors

View PDF HTML (experimental)

Abstract:Consensus among LLM judges is often taken as strong evidence that a decision is correct. This assumes that judges make their errors independently. In practice, LLM judges are often trained and evaluated in similar ways, so they can make the same mistakes. We study how this dependency affects the reliability of consensus. We find substantial error correlation across both open-weight and frontier LLM judges. In our main bank of ten judges, the average pairwise correlation between judge errors is 0.21. As a result, the ten judges only provide roughly as much statistical information as 3.5 independent judges. The dependency is even stronger among the high-accuracy frontier judges we evaluate, including judges from different providers. In up to 28% of our comparisons, ignoring shared errors leads to the conclusion that one system is significantly better, while accounting for them does not. We also find that the pattern of errors matters. Errors shared by most judges and errors concentrated among a smaller group affect consensus differently and favor different voting methods. Measuring the overall amount of correlation alone is therefore insufficient. Our results suggest a simple approach: use a small set of trusted examples to estimate judge accuracy and identify shared mistakes. These shared errors should then be considered when analyzing the results, and the voting method should be chosen using trusted examples before it is applied to new data.

Subjects:

Artificial Intelligence (cs.AI)

Cite as: arXiv:2609.22512 [cs.AI]

(or arXiv:2609.22512v1 [cs.AI] for this version)

https://doi.org/10.48550/arXiv.2609.22512

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Md Elias Hossain [view email] [v1] Fri, 18 Sep 2026 19:20:40 UTC (2,328 KB)

Full-text links:

Access Paper:

View a PDF of the paper titled Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus, by Elias Hossain and 2 other authors

View PDF

HTML (experimental)

TeX Source

view license

Current browse context:

cs.AI

new | recent | 2026-09

Change to browse by:

cs

References & Citations

NASA ADS

Google Scholar

Semantic Scholar

Loading...

Data provided by:

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

Author

Venue

Institution

Topic

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • LLM judges show substantial error correlation across open-weight and frontier models.
  • Ten judges with average pairwise error correlation 0.21 provide information equivalent to about 3.5 independent judges.
  • Shared errors changed significance in up to 28% of comparisons, and error patterns affect which voting methods work best.
  • The authors recommend using trusted examples to estimate accuracy, identify shared mistakes, and choose voting methods before deployment.

Highlights and analysis are generated automatically and may contain errors. Check the original source.