Skip to content
AI News HubLIVE
In-site rewrite3 min read

New framework helps robots turn complex language into precise 3D actions

Summary

Researchers at the Chinese University of Hong Kong and other institutes have developed Retrieval-Augmented Manipulation (RAM), a framework combining vision-language models with 3D object representations. It enables robots to interpret complex spatial instructions and execute tasks without task-specific training, demonstrated through zero-shot real-world tests. RAM adaptively replans actions based on physical constraints.

New framework helps robots turn complex language into precise 3D actions
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

May 22, 2026

feature

New framework helps robots turn complex language into precise 3D actions

by Ingrid Fadelli, Phys.org

Ingrid Fadelli

Author

edited by Stephanie Baum, reviewed by Robert Egan

Stephanie Baum

Scientific Editor

Robert Egan

Associate Editor

Editors' notes

This article has been reviewed according to Science X's editorial process and policies. Editors have highlighted the following attributes while ensuring the content's credibility:

fact-checked

peer-reviewed publication

trusted source

proofread

The GIST

Add as preferred source

Credit: Aideal Hwa, Unsplash.com

Over the past few decades, roboticists worldwide have introduced increasingly advanced robots that can understand human instructions, move in their surroundings and reliably complete basic manual tasks. While they perform well in some scenarios, many of these robots still struggle to translate the instructions of users into precise and executable actions that would allow them to successfully complete desired tasks.

Recently, computer scientists have been trying to improve how robots respond to user commands or queries using vision-language models (VLMs), artificial intelligence (AI) systems trained to process both images and texts. These models can typically interpret basic requests such as "place the bottle onto the plate," yet they often do not exhibit the spatial reasoning capabilities required to interpret more elaborate instructions and translate them into executable actions in real-world settings.

Researchers at the Chinese University of Hong Kong, the Zhejiang Humanoid Robot Innovation Center Co. Ltd and other institutes recently introduced Retrieval-Augmented Manipulation (RAM), a framework that could improve the ability of robots to connect abstract instructions with three-dimensional (3D) representations of the space around them. The new framework, presented in a Science Robotics paper, was found to improve the spatial reasoning capabilities of robots, allowing them to reliably follow more elaborate instructions, without requiring task-specific training.

"Although VLMs can interpret high-level commands, they lack the intrinsic spatial intelligence required for tasks demanding precise object placement, orientation, and physical reasoning," wrote Kai Chen, Chengkun Li and their colleagues in their paper. "We introduce Retrieval-Augmented Manipulation (RAM), an object-centric framework that endows general-purpose vision foundation models with the spatial reasoning necessary for robust manipulation."

The Retrieval-Augmented Manipulation (RAM) framework

The robotics framework developed by the researchers combines VLMs with explicit 3D object representations. In contrast with many previously proposed approaches, it acts as a bridge between two different capabilities, interpreting human instructions and making sense of how objects exist in 3D space.

"RAM bridges the semantic-to-geometric gap by grounding abstract concepts into an explicit, object-centric 3D representation," wrote the researchers. "This grounded information is then provided as augmented context to the VLM, empowering it to decompose complex instructions into a sequence of spatially precise and physically plausible subgoals."

Essentially, the RAM system analyzes images captured by a robot's integrated cameras, identifying specific objects and building a 3D object-centered representation of the current environment. This allows the model to delineate where objects are located, their approximate shapes/sizes, their orientations and how close they are to each other.

After a VLM processes the instructions provided by human users, the team's framework feeds spatial information from the 3D scene representations back to the model. This allows it to convert abstract language into instructions that are physically relevant to the present scenario.

The framework then breaks the task that the robot was instructed to complete into spatially informed subgoals. Breaking the tasks into smaller steps allows the system to adapt and plan different actions if something goes wrong or changes in the surrounding environment.

"We demonstrate that RAM, in a zero-shot setting on a real-world robot, can execute these subgoals to fulfill complex spatial language instructions, complete spatially aware manipulation under the guidance of a single 2D image, and adaptively replan tasks by reasoning about physical constraints like object size and collisions," wrote the authors. "Quantitative evaluations on the Common Object in 3D (CO3D) dataset also validated that RAM's core vision module generalizes to previously unseen object categories and is robust to variations in shape and occlusions."

Guiding the development of reliable and adaptable robots

The team had already tested their framework on a real robot, which was instructed to perform various tasks on which it had not been previously trained. Remarkably, they found that the robot could successfully complete many of these tasks, adaptively replanning its actions when a movement did not allow it to complete the desired sub-goals.

"By providing a structured bridge between semantic intent and geometric execution, RAM represents a critical step toward developing more physically intelligent and general-purpose robotic systems," wrote the researchers.

The team's framework could soon be refined further and tested in a wider range of real-world experiments with different robots, objects and user instructions. In the future, it could contribute to the advancement of household, industrial and service robots, allowing them to closely follow user instructions and flexibly adapt their actions in dynamic real-world environments.

Written for you by our author Ingrid Fadelli, edited by Stephanie Baum, and fact-checked and reviewed by Robert Egan—this article is the result of careful human work. We rely on readers like you to keep independent science journalism alive. If this reporting matters to you, please consider a donation (especially monthly). You'll get an ad-free account as a thank-you.

Publication details

Kai Chen et al, A retrieval-augmented framework enabling VLM spatial awareness for object-centric robot manipulation, Science Robotics (2026). DOI: 10.1126/scirobotics.aea2092

Journal information: Science Robotics

Key concepts

Embodied robotic manipulationHuman-centered AI interfacesComputational 3D visionAI wearablesAutonomous robotic locomotion

© 2026 Science X Network

Citation: New framework helps robots turn complex language into precise 3D actions (2026, May 22) retrieved 22 May 2026 from https://techxplore.com/news/2026-05-framework-robots-complex-language-precise.html

This document is subject to copyright. Apart from any fair dealing for the purpose of private study or research, no part may be reproduced without the written permission. The content is provided for information purposes only.

Explore further

Combining the robot operating system with LLMs for natural-language control

4 shares

Facebook

Twitter

Email

Feedback to editors

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • RAM bridges the semantic-to-geometric gap by grounding abstract concepts in explicit 3D object representations.
  • The framework decomposes complex instructions into spatially precise subgoals, allowing adaptive replanning.
  • Zero-shot tests on a real robot showed successful execution of novel tasks, accounting for object size and collisions.
  • Potential applications include household, industrial, and service robots requiring flexible spatial reasoning.

Highlights and analysis are generated automatically and may contain errors. Check the original source.