PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
PerceptionBench is a benchmark designed to evaluate atomic visual perception in multimodal large language models (MLLMs). It uses a bottom-up approach, building an error taxonomy from 42 existing benchmarks to define ten atomic perceptual capabilities, and creates 3,000 verified questions. Tests on 16 frontier MLLMs show no model exceeds 60% accuracy, with perception-related hallucination being the weakest capability.
-->
[Submitted on 27 Jul 2026]
Title:PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
View a PDF of the paper titled PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models, by Zichao Lin and 32 other authors
View PDF HTML (experimental)
Abstract:We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60\% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.
Subjects:
Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2607.24957 [cs.CV]
(or arXiv:2607.24957v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2607.24957
arXiv-issued DOI via DataCite (pending registration)
Submission history
From: Bowen Qu [view email] [v1] Mon, 27 Jul 2026 18:04:54 UTC (10,386 KB)
Full-text links:
Access Paper:
View a PDF of the paper titled PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models, by Zichao Lin and 32 other authors
View PDF
HTML (experimental)
TeX Source
view license
Current browse context:
cs.CV
new | recent | 2026-07
Change to browse by:
cs
References & Citations
NASA ADS
Google Scholar
Semantic Scholar
Loading...
Data provided by:
Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Bibliographic Explorer (What is the Explorer?)
Connected Papers Toggle
Connected Papers (What is Connected Papers?)
Litmaps Toggle
Litmaps (What is Litmaps?)
scite.ai Toggle
scite Smart Citations (What are Smart Citations?)
Code, Data, Media
Code, Data and Media Associated with this Article
alphaXiv Toggle
alphaXiv (What is alphaXiv?)
Links to Code Toggle
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub Toggle
DagsHub (What is DagsHub?)
GotitPub Toggle
Gotit.pub (What is GotitPub?)
Huggingface Toggle
Hugging Face (What is Huggingface?)
ScienceCast Toggle
ScienceCast (What is ScienceCast?)
Demos
Demos
Replicate Toggle
Replicate (What is Replicate?)
Spaces Toggle
Hugging Face Spaces (What is Spaces?)
Spaces Toggle
TXYZ.AI (What is TXYZ.AI?)
Related Papers
Recommenders and Search Tools
Link to Influence Flower
Influence Flower (What are Influence Flowers?)
Core recommender toggle
CORE Recommender (What is CORE?)
Author
Venue
Institution
Topic
About arXivLabs
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)