Skip to content
AI News HubLIVE
Original source2 min read

The Wisdom of Artificial Deliberative Crowds

Summary

A new arXiv paper adapts a three-stage human deliberation paradigm to large language models from three different families and tests it across four domains of increasing real-world stakes. Deliberation reduced collective error beyond passive aggregation of independent answers, and post-deliberation individual judgments retained that collective gain—but the advantage required model diversity, as groups composed of clones from a single model did not benefit.

SourcearXiv AIAuthor: Federico Barrera-Lemarchand, Mariano Sigman, Joaquin Navajas
The Wisdom of Artificial Deliberative Crowds
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

[Submitted on 18 Sep 2026]

Title:The Wisdom of Artificial Deliberative Crowds

View a PDF of the paper titled The Wisdom of Artificial Deliberative Crowds, by Federico Barrera-Lemarchand and 2 other authors

View PDF

Abstract:The aggregation of many lay estimates often outperforms individual expert judgment, a phenomenon known as the wisdom of crowds. While this is usually attributed to the independence of estimates, an even stronger effect arises through deliberation: averaging the consensus estimates of small deliberating groups outperforms the classical wisdom of crowds, with individual judgments themselves also becoming more accurate after deliberation. Whether these improvements transfer to large language models deliberating amongst themselves is unknown. Here we adapt a three-stage deliberation paradigm previously used with human participants for use with large language models from three different families, and test it across four domains of increasing real-world stakes: visual numerical estimation (Study 1), peer review of machine-learning papers (Study 2), detection of hidden malicious behavior by an artificial intelligence agent (Study 3), and sports forecasting against a real prediction market (Study 4). Across domains, deliberation reduced collective error beyond passive aggregation of independent responses, and post-deliberation individual judgments retained this collective gain. Notably, the advantage required model diversity: groups composed of clones of a single model did not benefit from deliberating. These results establish machine deliberation as a general-purpose aggregation mechanism, and point to diversity as an active ingredient.

Subjects:

Artificial Intelligence (cs.AI)

Cite as: arXiv:2609.22497 [cs.AI]

(or arXiv:2609.22497v1 [cs.AI] for this version)

https://doi.org/10.48550/arXiv.2609.22497

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Joaquin Navajas [view email] [v1] Fri, 18 Sep 2026 19:01:43 UTC (1,323 KB)

Full-text links:

Access Paper:

View a PDF of the paper titled The Wisdom of Artificial Deliberative Crowds, by Federico Barrera-Lemarchand and 2 other authors

View PDF

view license

Current browse context:

cs.AI

new | recent | 2026-09

Change to browse by:

cs

References & Citations

NASA ADS

Google Scholar

Semantic Scholar

Loading...

Data provided by:

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

Author

Venue

Institution

Topic

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • The study adapts a three-stage deliberation paradigm previously used with humans for LLMs from three different model families.
  • Across visual numerical estimation, ML paper peer review, hidden malicious behavior detection, and sports forecasting, deliberation cut collective error beyond passive aggregation.
  • Model diversity is essential: groups of clones from a single model gained nothing from deliberating.

Highlights and analysis are generated automatically and may contain errors. Check the original source.