DexAgent

An Agentic Human2Sim2Robot Framework for Dexterous Manipulation with Self-Evolving Tool Library

Code (Coming Soon!)
S C R O L L

Overview Video

Overview

DexAgent overview
DexAgent converts a single egocentric human video and a task prompt into robot training data through four stages, selecting property-specific skills and verifiers from a self-evolving tool library. We evaluate on eleven dexterous manipulation tasks spanning rigid, articulated, deformable and long-horizon interactions.

Abstract

Human videos offer a scalable source of demonstrations for dexterous robot manipulation. However, existing human-to-simulation-to-robot (Human2Sim2Robot) pipelines rely on predefined procedures that struggle to accommodate diverse object properties and interactions, particularly those involving articulated and deformable objects. We introduce DexAgent, an agentic Human2Sim2Robot framework that converts a single egocentric human video and a task prompt into physically grounded robot trajectories for policy training. It operates through four stages: semantic understanding of human videos, property-based simulation reconstruction, robot trajectory optimization, and robot data generation. At each stage, DexAgent adapts its approach to the task and object properties by selecting suitable skills from its tool library or developing new ones when needed. Property-specific verifiers assess stage outcomes for physical validity and task-specific requirements and provide feedback for refinement, preventing error propagation through the workflow. In the final stage, DexAgent varies object and robot states in simulation to generate diverse robot trajectories from a single human video, then retextures the rendered observations to facilitate sim-to-real transfer. Newly developed skills and verifiers are retained in its tool library, making it self-evolving to accumulate reusable capabilities and reducing processing time as it encounters more human videos. Across eleven real-world tasks, policies trained with DexAgent-generated data achieve a 3.5× higher success rate than competing baselines.

Methodology

DexAgent method
Stage 1 decomposes the task into subgoals and identifies the manipulated objects and their key properties. Stage 2 reconstructs the objects in simulation using property-specific skills from the tool library, writing new ones when none fits. Stage 3 generates robot-hand trajectories by combining optimization guided by human motion priors with code-generated motion. Stage 4 varies viewpoints and object states to produce diverse simulation data, inpaints the original human video, and retextures the augmented renders for policy training.

Data Quality Experiments

Simulation Reconstruction

HOI4D

Method Rigid Articulated
F-5↑ F-10↑ CD↓ F-5↑ F-10↑ CD↓
HO0.280.513.860.290.471.30
IHOI0.420.702.700.320.471.47
HORSE0.260.456.690.190.341.91
MCC-HO0.520.781.360.350.551.21
G-HOP0.690.910.630.070.091.23
FoundationPose0.710.910.490.400.601.26
Any6D0.710.910.500.380.601.24
Do as I Do0.720.910.490.400.611.25
Ours 0.830.960.29 0.470.681.08
Reconstruction accuracy on HOI4D, scored separately for rigid and articulated objects. Evaluation metrics are F-5 and F-10, which are F-scores at two distance thresholds, and CD, which is Chamfer distance. Best entry in each column in bold.

Trajectory Retargeting

OakInk

Method Success %↑ Epos (m)↓ Erot (rad)↓
Dex-retargeting28.60.080.62
SPIDER mjwp71.40.040.57
SPIDER mjwp_act77.10.040.42
Do-as-I-Do Sharpa hand81.00.030.15
without transition reward79.00.030.14
annealed sampling only72.00.080.32
Ours 85.70.030.12
Human-to-robot retargeting quality on OakInk. Success is the share of clips whose converted trajectory passes the physics check. Epos and Erot are the residual position and rotation gap from the demonstration, lower is better. The two indented rows ablate the strongest baseline.

Real-Robot Experiments

Replay

MethodCupGiftboxDrawerRope KnotScissorsDrawingComputerBottleBiologyBatteryMulti-objectTotal
Dex-retargeting✓✗✗✗✗✗✗✗✗✗✗1/11
Do as I Do✓✗✗✗✗✗✗✗✗✗✗1/11
Spider✓✓✓✗✓✗✓✗✗✓✗6/11
TopoRetarget✓✓✓✗✗✗✗✗✗✗✗3/11
Egoinfinity✓✗✗✗✗✗✗✗✗✗✗1/11
V2D✓✓✓✗✓✗✗✗✗✗✓5/11
GPT-6: Astra✓✓✓✗✗✓✓✗✗✗✗5/11
Ours✓✓✓✓✓✓✓✓✓✓✓11/11
Real-world replay success across eleven tasks. Each converted trajectory is executed open loop, which scores the conversion rather than a learned policy. A task is marked successful if at least one of ten replay trials succeeds. GPT-6 Astra joined the benchmark after the first evaluation round.

Policy

0 20 40 60 80 100 Success Rate (%) 90 70 80 50 80 40 50 70 40 60 70 63.6 Cup Giftbox Drawer Rope Knot Scissors Drawing Computer Bottle Biology Battery Multi-object Average
Per-task policy success rate across eleven tasks. Policies trained on DexAgent-generated data reach 63.6% on average, against 18.2% for the strongest baseline.

A Self-Evolving Tool Library

Breakdown of the DexAgent tool library: 103 skills and 188 verifiers, grouped by functionality
Breakdown of the current DexAgent tool library. The library contains 103 skills and 188 verifiers grouped by functionality. Skills cover perception, reconstruction, grasp and pose generation, scene execution, measurement, and trajectory generation, while verifiers check task completion, grasp stability, reconstructed mechanisms, full-episode consistency, task-specific scene conditions, and scene validity. These skills and verifiers are expected to generalize to new and unseen samples, and the library is expected to continue to grow as DexAgent processes more new samples.
Processing cost falls as the library grows What drives each step in library growth
Processing cost falls as the library grows. Over 100 EgoDex samples the library reaches 85 skills and 168 verifiers. In terms of cost, V2D converting one sample takes 3.7 hours on average and GPT 6: Astra takes 3.3 hours on average at a flat rate, whereas DexAgent takes 2.1 hours on average in a fresh run, a 36.4% reduction comparing to GPT 6: Astra and a 43.2% reduction comparing to V2D.

Simulation and Rollout Videos

Simulation 6×

Human Demonstration

Open the Bottle

Rollout 6×

Open the Bottle