AI News HubLIVE
Original source2 min read

Bayesian Wind Tunnels for Model Selection

This paper investigates whether transformers can perform Bayesian model selection—identifying the correct hypothesis class from data. Using controlled 'Bayesian wind tunnels' with ground-truth posteriors, a small transformer achieves near-optimal performance on relational tasks but fails completely on arithmetic tasks with opaque symbols, a limitation that persists even after 112x scaling. Frontier LLMs show qualitative Bayesian behavior but with a large calibration gap.

SourcearXiv Machine LearningAuthor: Siddhartha R Dalal, Vishal Misra, Abhay Parekh

-->

[Submitted on 1 Jul 2026]

Title:Bayesian Wind Tunnels for Model Selection

View a PDF of the paper titled Bayesian Wind Tunnels for Model Selection, by Siddhartha R Dalal and Vishal Misra and Abhay Parekh

View PDF HTML (experimental)

Abstract:Prior work has shown that transformers can perform exact Bayesian filtering within a fixed

hypothesis class. Can they also perform Bayesian model selection -- identifying the correct

hypothesis class from data? We introduce model-selection Bayesian wind tunnels: controlled

environments where ground-truth posteriors over hypothesis classes are available in closed

form. Using fixed-point-free involutions -- whose defining property f(f(x))=x is purely

relational -- a 2.8M-parameter transformer achieves 0.01-bit entropy agreement with the

Bayesian optimum (3 seeds), with both integer tokens and opaque symbols whose meanings

change every episode. This extends to non-nested comparisons: involutions vs. 3-cycles

(where neither class is a subset of the other) achieve class-posterior MAE under 0.001,

demonstrating genuine model selection beyond simplicity/subset bias. We then identify a

sharp perceptual access condition: when the discriminative statistic requires arithmetic --

modular addition (rotations) or multiplication (f(x)=cx mod p) -- model selection succeeds

with integer tokens but fails completely with opaque symbols, and this boundary persists

under 112x scaling (2.8M to 316M parameters). A stationarity control confirms the operative

factor: opaque tokens with a fixed relabeling succeed (0.009-bit MAE), showing that stable

semantics, not integer identity, enable circuit compilation. Header subtask diagnostics

localize the failure to the composition of header inversion with arithmetic rather than

header parsing itself. Probing frontier LLMs on the same tasks shows qualitative Bayesian

behavior but a large calibration gap (~55x), measured through lossy probes and therefore

directional rather than exact.

Comments: 27 pages

Subjects:

Machine Learning (cs.LG); Machine Learning (stat.ML)

ACM classes: I.2.6; G.3

Cite as: arXiv:2607.19379 [cs.LG]

(or arXiv:2607.19379v1 [cs.LG] for this version)

https://doi.org/10.48550/arXiv.2607.19379

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Vishal Misra [view email] [v1] Wed, 1 Jul 2026 04:05:08 UTC (124 KB)

Full-text links:

Access Paper:

View a PDF of the paper titled Bayesian Wind Tunnels for Model Selection, by Siddhartha R Dalal and Vishal Misra and Abhay Parekh

View PDF

HTML (experimental)

TeX Source

view license

Current browse context:

cs.LG

new | recent | 2026-07

Change to browse by:

cs stat stat.ML

References & Citations

NASA ADS

Google Scholar

Semantic Scholar

Loading...

Data provided by:

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

IArxiv recommender toggle

IArxiv Recommender (What is IArxiv?)

Author

Venue

Institution

Topic

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)