May 22, 2026
feature
New framework helps robots turn complex language into precise 3D actions
by Ingrid Fadelli, Phys.org
Ingrid Fadelli
Author
edited by Stephanie Baum, reviewed by Robert Egan
Stephanie Baum
Scientific Editor
Robert Egan
Associate Editor
Editors' notes
This article has been reviewed according to Science X's editorial process and policies. Editors have highlighted the following attributes while ensuring the content's credibility:
fact-checked
peer-reviewed publication
trusted source
proofread
The GIST
Add as preferred source
Credit: Aideal Hwa, Unsplash.com
Over the past few decades, roboticists worldwide have introduced increasingly advanced robots that can understand human instructions, move in their surroundings and reliably complete basic manual tasks. While they perform well in some scenarios, many of these robots still struggle to translate the instructions of users into precise and executable actions that would allow them to successfully complete desired tasks.
Recently, computer scientists have been trying to improve how robots respond to user commands or queries using vision-language models (VLMs), artificial intelligence (AI) systems trained to process both images and texts. These models can typically interpret basic requests such as "place the bottle onto the plate," yet they often do not exhibit the spatial reasoning capabilities required to interpret more elaborate instructions and translate them into executable actions in real-world settings.
Researchers at the Chinese University of Hong Kong, the Zhejiang Humanoid Robot Innovation Center Co. Ltd and other institutes recently introduced Retrieval-Augmented Manipulation (RAM), a framework that could improve the ability of robots to connect abstract instructions with three-dimensional (3D) representations of the space around them. The new framework, presented in a Science Robotics paper, was found to improve the spatial reasoning capabilities of robots, allowing them to reliably follow more elaborate instructions, without requiring task-specific training.
"Although VLMs can interpret high-level commands, they lack the intrinsic spatial intelligence required for tasks demanding precise object placement, orientation, and physical reasoning," wrote Kai Chen, Chengkun Li and their colleagues in their paper. "We introduce Retrieval-Augmented Manipulation (RAM), an object-centric framework that endows general-purpose vision foundation models with the spatial reasoning necessary for robust manipulation."
The Retrieval-Augmented Manipulation (RAM) framework
The robotics framework developed by the researchers combines VLMs with explicit 3D object representations. In contrast with many previously proposed approaches, it acts as a bridge between two different capabilities, interpreting human instructions and making sense of how objects exist in 3D space.
"RAM bridges the semantic-to-geometric gap by grounding abstract concepts into an explicit, object-centric 3D representation," wrote the researchers. "This grounded information is then provided as augmented context to the VLM, empowering it to decompose complex instructions into a sequence of spatially precise and physically plausible subgoals."
Essentially, the RAM system analyzes images captured by a robot's integrated cameras, identifying specific objects and building a 3D object-centered representation of the current environment. This allows the model to delineate where objects are located, their approximate shapes/sizes, their orientations and how close they are to each other.
After a VLM processes the instructions provided by human users, the team's framework feeds spatial information from the 3D scene representations back to the model. This allows it to convert abstract language into instructions that are physically relevant to the present scenario.
The framework then breaks the task that the robot was instructed to complete into spatially informed subgoals. Breaking the tasks into smaller steps allows the system to adapt and plan different actions if something goes wrong or changes in the surrounding environment.
"We demonstrate that RAM, in a zero-shot setting on a real-world robot, can execute these subgoals to fulfill complex spatial language instructions, complete spatially aware manipulation under the guidance of a single 2D image, and adaptively replan tasks by reasoning about physical constraints like object size and collisions," wrote the authors. "Quantitative evaluations on the Common Object in 3D (CO3D) dataset also validated that RAM's core vision module generalizes to previously unseen object categories and is robust to variations in shape and occlusions."
Guiding the development of reliable and adaptable robots
The team had already tested their framework on a real robot, which was instructed to perform various tasks on which it had not been previously trained. Remarkably, they found that the robot could successfully complete many of these tasks, adaptively replanning its actions when a movement did not allow it to complete the desired sub-goals.
"By providing a structured bridge between semantic intent and geometric execution, RAM represents a critical step toward developing more physically intelligent and general-purpose robotic systems," wrote the researchers.
The team's framework could soon be refined further and tested in a wider range of real-world experiments with different robots, objects and user instructions. In the future, it could contribute to the advancement of household, industrial and service robots, allowing them to closely follow user instructions and flexibly adapt their actions in dynamic real-world environments.
Written for you by our author Ingrid Fadelli, edited by Stephanie Baum, and fact-checked and reviewed by Robert Egan—this article is the result of careful human work. We rely on readers like you to keep independent science journalism alive. If this reporting matters to you, please consider a donation (especially monthly). You'll get an ad-free account as a thank-you.
Publication details
Kai Chen et al, A retrieval-augmented framework enabling VLM spatial awareness for object-centric robot manipulation, Science Robotics (2026). DOI: 10.1126/scirobotics.aea2092
Journal information: Science Robotics
Key concepts
Embodied robotic manipulationHuman-centered AI interfacesComputational 3D visionAI wearablesAutonomous robotic locomotion
© 2026 Science X Network
Citation: New framework helps robots turn complex language into precise 3D actions (2026, May 22) retrieved 22 May 2026 from https://techxplore.com/news/2026-05-framework-robots-complex-language-precise.html
This document is subject to copyright. Apart from any fair dealing for the purpose of private study or research, no part may be reproduced without the written permission. The content is provided for information purposes only.
Explore further
Combining the robot operating system with LLMs for natural-language control
4 shares
Feedback to editors