Skip to content
AI News HubLIVE
Original source2 min read

FailureSpot: Label-Efficient Timestamp-Level Failure Detection for Vision-Language-Action Models

Summary

Vision-language-action (VLA) policies can fail unpredictably in long-horizon tasks, and existing detection approaches often suffer from trajectory-level label noise. FailureSpot uses unlabeled action chunks to derive weak supervision signals and then applies active learning to select the most uncertain trajectories for timestamp-level annotation, improving both timestamp-level and trajectory-level failure detection across multiple VLA policies.

SourcearXiv RoboticsAuthor: Jie Ma, Zongxi Liu, Yi Zhu
FailureSpot: Label-Efficient Timestamp-Level Failure Detection for Vision-Language-Action Models
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

[Submitted on 3 Sep 2026]

Title:FailureSpot: Label-Efficient Timestamp-Level Failure Detection for Vision-Language-Action Models

View a PDF of the paper titled FailureSpot: Label-Efficient Timestamp-Level Failure Detection for Vision-Language-Action Models, by Jie Ma and 2 other authors

View PDF HTML (experimental)

Abstract:Vision-language-action (VLA) policies have shown strong potential for general-purpose robotic manipulation, but they can still fail unpredictably during long-horizon execution, making reliable failure detection essential for safe deployment. Existing methods either rely on visual models that typically detect failures only after erroneous actions have occurred, or use lightweight proactive detectors trained on VLA internal representations. However, these proactive methods are often supervised with trajectory-level labels, causing normal pre-failure behavior in unsuccessful trajectories to be incorrectly labeled as failure. This supervision mismatch introduces label noise and limits both trajectory-level detection accuracy and precise timestamp-level failure localization. In this work, we study fine-grained timestamp-level VLA failure detection while addressing the cost of dense annotation. We propose a data-efficient framework that first leverages unlabeled VLA action chunks to construct action-derived weak supervision signals, capturing abnormal patterns such as inconsistent consecutive chunks, frozen or idle actions, and aggressive random motions. We then use active learning to select only the most uncertain trajectories for timestamp-level annotation and fine-tune the detector with these informative labels. Experiments across multiple VLA policies show that our method improves both timestamp-level and trajectory-level failure detection performance.

Subjects:

Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

Cite as: arXiv:2609.04277 [cs.RO]

(or arXiv:2609.04277v1 [cs.RO] for this version)

https://doi.org/10.48550/arXiv.2609.04277

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yi Zhu [view email] [v1] Thu, 3 Sep 2026 00:04:06 UTC (4,765 KB)

Full-text links:

Access Paper:

View a PDF of the paper titled FailureSpot: Label-Efficient Timestamp-Level Failure Detection for Vision-Language-Action Models, by Jie Ma and 2 other authors

View PDF

HTML (experimental)

TeX Source

view license

Current browse context:

cs.RO

new | recent | 2026-09

Change to browse by:

cs cs.CV

References & Citations

NASA ADS

Google Scholar

Semantic Scholar

Loading...

Data provided by:

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

Author

Venue

Institution

Topic

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • Existing visual failure detectors typically react after erroneous actions, while proactive detectors trained on VLA representations are misled by trajectory-level labels.
  • FailureSpot builds weak supervision from action chunks, identifying abnormal patterns like inconsistent chunks, frozen actions, and aggressive random motions.
  • Active learning selects the most uncertain trajectories for timestamp-level labeling, cutting annotation cost while boosting detection precision.

Highlights and analysis are generated automatically and may contain errors. Check the original source.