UMR: Universal Manipulation Representation

Song Liu1,2,*, Linyi Li1,2,*
Yanshun Zhao1, Rxuan Li1, Xinrui Xu1, Yi Ju1, Yahui Deng1, Senge Zhang1, Guoyu Liu1, Yixuan Li1,
Wuyang Zhang1,2, Yao Li1, Congcong Zhu1,2,†, Jingrun Chen1,†
1University of Science and Technology of China, Hefei, China.
2Suzhou Artificial Intelligence Laboratory, Suzhou, China.
* Equal contribution. † Corresponding authors.

Replace this sentence with your one-line paper summary.

Abstract

UMR decomposes manipulation into World Flow and Ego Trajectory, coupling transferable object motion with executable robot-local actions.

Method Overview

Universal Manipulation Representation

Two coordinate views of the same underlying physical motion.

World Flow describes task-relevant object motion as a structured SE(3) trajectory in the world frame. Ego Trajectory expresses the corresponding end-effector motion relative to its current pose, yielding a controller-aligned action representation.

World FlowTransferable task motion
SE(3)conjugation
Ego TrajectoryExecutable local motion

The conjugate transformation makes them geometrically consistent views of one motion, allowing task intent learned from one demonstrator to be executed by another embodiment.

UMR representation overview · excerpt from 00:20–00:37

Experiment

Settings

Simulation

Benchmark Task Settings

Overview of the ten RLBench manipulation tasks
RLBench10 manipulation tasks
Overview of the four LIBERO task suites
LIBERO4 task suites · 40 tasks
Real World

Cross-Condition Settings

UMR real-world robot setups and six evaluation settings
Real-World SettingHuman-to-robot transfer without robot demonstrations
2Robot
embodiments
2Tabletop
scenes
3Gripper
configurations
6Evaluation
settings
Simulation

Task Demonstrations

RLBench · Close Box Articulated-object manipulation
Real World

Task Demonstrations

Cube Stacking Real-world human demonstration for cube stacking

Results

Simulation Benchmarks

WEPVLA instantiates UMR as a compact 0.5B-parameter point-cloud VLA. Both World Flow and Ego Trajectory are jointly denoised at inference, while only the locally executable Ego action chunk is sent to the robot.

RLBench 85.7%

Mean success across 10 tasks, outperforming PointACT (82.3%) while using a smaller 0.5B model.

LIBERO 97.5%

Average success across Spatial, Object, Goal, and Long-Horizon suites with one policy trained on all 40 tasks.

Representation World + Ego

World Flow captures transferable task motion; Ego Trajectory provides executable robot-local control.

Benchmark Model Model Size Mean / Avg. Success
RLBench 10 tasksHybridVLA7B74.0
RLBench 10 tasksPointACT3B82.3
RLBench 10 tasksWEPVLA (Ours)0.5B85.7
LIBERO 4 suitesPointACT3B96.0
LIBERO 4 suitesOpenVLA-OFT7B96.8
LIBERO 4 suitesWEPVLA (Ours)0.5B97.5

Real-World Transfer

Method Training Data Cross-Environment Cross-Embodiment Cross-Setup Tasks I–VI Avg.
HumanEgoHuman only, ≈10 min/task70.060.858.360.8
WEPVLA (Ours)Human only, ≈10 min/task95.091.790.091.7
Real-World Transfer 91.7%

Average across six settings with no robot demonstrations, versus 60.8% for HumanEgo.

Human Data ≈10 min

Per task: 100 human demonstrations, with DES generating 500 additional episodes.

Data Scaling 82% → 96%

Long-6/Long-8 average when scaling from 100 to 14,300 heterogeneous episodes.

Heterogeneous Demonstration Gallery

The V10 scaling study combines heterogeneous demonstrations from RH20T, LIBERO, and RLBench with egocentric and exocentric real-world observations. The Data-Efficient Strategy further expands task-relevant object configurations through stage-aware point-cloud editing while preserving interaction geometry.

RH20T

Real RGB-D manipulation demonstrations, including pressing, napkin pulling, and drawer opening.

LIBERO

Spatial, object, goal, and long-horizon tasks reconstructed as RGB-D demonstrations.

RLBench

Simulation skills spanning articulated objects, grasp manipulation, tool use, and watering.

Egocentric

First-person human manipulation demonstrations captured from a body-aligned viewpoint.

Exocentric

Third-person observations of the same manipulation setting from an external camera viewpoint.

Data Efficient Strategy

Stage-aware point-cloud editing generates diverse object configurations while preserving interaction geometry.

Interactive Results

Pick a scene below to view RGB-D inputs, prediction, and ground truth.

  • Policy: WEPVLA learns coupled World Flow and Ego Trajectory action chunks.
  • Inputs: RH20T RGB-D observations and virtual gripper trajectories.
  • Outputs: single-view filtered point clouds aligned with multiview references.
UMR RH20T

NOTE: Prediction shows single-view filtering; the reference viewer shows multiview reconstruction.

RH20T Frame and Review Projection Hover to zoom

RGB cam0
Projection comparison
Single vs multiview
Source RGB

Single-View Prediction

choose a scene above to view interactive visualization

Multiview Reference

choose a scene above to view interactive visualization

Evaluation Videos

FrankaCube Stacking
Rollout 1 of 5

Select a task and rollout to inspect UMR under different embodiments and real-world manipulation conditions.

Failure Cases

Select a case to inspect the rollout and the mechanism behind the observed failure.

Acknowledgements

We thank the colleagues who provided valuable feedback on the UMR formulation and assisted with data collection, system integration, and real-world robot evaluations. We are also grateful to the creators and maintainers of RLBench, LIBERO, RH20T, SmolVLA, and Viser, as well as the broader open-source robotics community, for the benchmarks, datasets, model components, and visualization tools that supported this work.