Skip to content
AI News HubLIVE
Original source2 min read

Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes

Summary

A new arXiv paper introduces Segment-Snap, a framework that couples part segmentation, motion constraints, and operable-region prediction through the physical relationship between parts and handles. On Articulate3D validation, handle guidance lifts motion-gated AP from 13.74% to 40.98%, and full context reaches 30.99% handle AP.

SourcearXiv Computer VisionAuthor: Hanyang Kong, Xingyi Yang
Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

[Submitted on 21 Sep 2026]

Title:Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes

View a PDF of the paper titled Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes, by Hanyang Kong and 1 other authors

View PDF HTML (experimental)

Abstract:Interaction understanding in 3D scenes requires a joint description of movable parts, their motion, and the regions through which they can be operated. We present Segment-Snap, which connects these outputs through the physical relationship between parts and handles. Learned predictors identify broad part surfaces and small handles. A geometric decoder uses planar and upright priors to constrain motion, then selects hinge lines using predicted handle locations, without training a motion regressor. Conversely, a joint part-and-handle predictor supplies additional handle candidates, whose motion classes are refined using containing parts. Each information transfer is applied once, without iterative feedback. On Articulate3D validation, handle guidance raises motion-gated AP from 13.74% to 40.98% at fixed masks and axes. Additional handle candidates raise handle AP from 24.63% to 29.65%; part-based class correction adds 0.98 points, and full context reaches 30.99%. Repeated training, learned-decoder controls and paired visualizations establish the benefits and limitations of combining geometric and semantic evidence for interaction understanding.

Comments: Project page: this https URL

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

Cite as: arXiv:2609.25247 [cs.CV]

(or arXiv:2609.25247v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.25247

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Hanyang Kong [view email] [v1] Mon, 21 Sep 2026 18:02:34 UTC (2,201 KB)

Full-text links:

Access Paper:

View a PDF of the paper titled Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes, by Hanyang Kong and 1 other authors

View PDF

HTML (experimental)

TeX Source

view license

Current browse context:

cs.CV

new | recent | 2026-09

Change to browse by:

cs cs.AI

References & Citations

NASA ADS

Google Scholar

Semantic Scholar

Loading...

Data provided by:

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

Author

Venue

Institution

Topic

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Key points and analysis

Article intelligence

ResearchersAdvanced

Key points

  • Segment-Snap links movable parts, their motion, and operable regions via the physical relationship between parts and handles.
  • A geometric decoder uses planar and upright priors to constrain motion and picks hinge lines from predicted handle locations, without training a motion regressor.
  • On Articulate3D, handle guidance raises motion-gated AP from 13.74% to 40.98%, while full context reaches 30.99% handle AP.
  • Information transfers are single-pass with no iterative feedback; repeated training, learned-decoder controls, and paired visualizations map both the benefits and limits of combining geometric and semantic evidence.

Highlights and analysis are generated automatically and may contain errors. Check the original source.