IronMind

Scaling Humanoid Dexterous Manipulation via Camera-Space Ego-Centric Pretraining

XPENG Robotics Contributors

IronMind on IRON-R01

Fully autonomous · all videos play at 1× speed

Overview

Turning egocentric human video into robot manipulation knowledge

Egocentric human video is an extraordinarily scalable data source for dexterous manipulation, but two things block it: an embodiment gap — human hands are not robot end-effectors, and cheap recordings carry no torso kinematics for retargeting — and heterogeneous quality, from noisy hand tracking to weakly aligned text.

IronMind bypasses body retargeting by supervising actions directly in camera space, curates 10,000+ hours through filtering, re-annotation and per-frame weighting, and learns with a Mixture-of-Transformers policy coupling vision-language understanding to a flow-matching action expert. Validation loss falls log-linearly with data scale, and the gains carry to out-of-distribution real-robot manipulation.

Data

10,000+ hours, all in one camera frame

Camera space as the action interface

Mapping human motion into a robot torso frame requires assumptions about a body the headset never sees — a rigid, vertical spine — and every assumption becomes label noise. IronMind keeps human trajectories in the native camera frame, transforms robot trajectories into the head camera’s frame by calibration, and aligns the two embodiments by meaning: wrist pose and finger flexion share vector dimensions, so human motion supervises the outputs the robot actually uses.

Donut chart of the pretraining corpus by scene category, with representative frames from each.
Figure 2Scene distribution of the pretraining corpus — pick a category to see frames from it.

Curated, not just collected

Raw recordings are full of passive observation, contradictory captions and task labels too coarse to separate sub-actions. Four stages fix that: rule-based filtering drops unrecoverable intrinsics and implausible hand poses; atomic re-annotation cuts recordings into single-intent segments, re-captioned by a VLM, with action losses masked at the boundaries; velocity alignment slows agile human motion toward teleoperation speed; and per-frame quality weighting scores every frame, letting the dataloader sample smooth, intentful motion — 51% effective sample size, no dataset regeneration.

The three curation stages and the data retention funnel.
Figure 3Rule-based filtering, atomic task re-annotation and per-frame quality weighting, with the retention funnel on the right.
Architecture

VLA with World Priors

A vision-language understanding expert and a flow-matching action expert, joined by shared attention at selected depths.

IronMind architecture: inputs, Mixture-of-Transformers, and train-only prior branches.

Three auxiliary heads — future frames, metric depth, semantic features — hang off a small set of learnable query tokens. Every teacher stays frozen and gradients reach the backbone only through those tokens. After pretraining the heads are deleted: the queries remain, and inference latency is unchanged.

Open-loop camera-space Evaluation

Scoring a policy without deploying on hardware

Evaluation protocol: open-loop metrics, the evaluation loop used during pretraining, and in-the-loop rollout.

Open-loop evaluation requires only an image (with estimated depth information) with hands and objects. We run the model on the image for one step, producing one action chunk, and evaluate the quality of that action chunk. For example, if the prompt is to grab an object in the image, then the action chunk should move the hand closer to that object. We calculate the distance between hand and object closed by the action chunk produced by the policy as an evaluation metric. If the hand is close enough to the object, we also calculate a grasping metric, i.e. does the action chunk produce a grasping movement.

Full set of open-loop evaluation metrics are:

  • dₕ₁₀ 10th percentile of hand-to-target distance over the chunk, cm — closest approach, stray steps discounted. Lower is better.
  • progress share of the initial gap closed by the last step. Higher is better.
  • θalign angle between the direction to the target and the wrist’s net displacement. Lower is better.
  • eFA mean absolute finger-aperture error against the demonstration, mm. This measures grasping accuracy. Lower is better.

We carefully construct our open-loop evaluation set to include various OOD scenarios:

Predicted hand trajectories rendered on unseen scenes, grouped by semantic, affordance and fine-grained manipulation.
Figure 6Open-loop evaluations probe unseen scenes: what to interact with, where, and how. Grasp aperture and wrist orientation adapt to the target’s geometry.
Scaling

Validation loss falls log-linearly, with no saturation at 10,000 hours

Validation loss against training step for each pretraining budget, and best validation loss against pretraining hours on a log scale.
Figure 7Pretraining data scaling. The 250-hour and 500-hour budgets overfit; from 2,000 hours upward the best validation loss tracks a straight line in log hours.

The gain is not confined to validation loss. Intermediate budgets are not strictly ordered on every metric, but target approach improves monotonically with data and the largest budget is best in all four columns.

Table 1. Open-loop performance vs. pretraining budget (one chunk).
Budgetdₕ₁₀ (cm) ↓Prog. (%) ↑θalign (°) ↓eFA (mm) ↓
250 h37.814.747.819.97
500 h33.428.141.719.54
2,000 h32.030.736.023.83
5,000 h30.834.041.220.33
10,000 h25.844.926.015.66

Chaining three chunks preserves the same ranking.

Table 2. Policy-in-the-loop performance vs. pretraining budget (three chunks, 970 cases).
Budgetdₕ₁₀ (cm) ↓Prog. (%) ↑θalign (°) ↓
250 h32.021.446.1
500 h25.632.239.5
2,000 h23.039.033.9
5,000 h22.739.136.3
10,000 h14.649.321.6

Raters preferred the 10,000-hour checkpoint in 47 of the 70 decided votes.

Table 3. Human preference, 20 held-out trials, 100 votes — 30 ties, 70 decided.
Pretraining budgetVotesWin rate ↑
250 h2332.9%
10,000 h4767.1%
Ablation

What the world-model priors buy

All arms share one pretraining recipe and differ from the three-prior arm by a single key, scored on the same 970-case approaching set.

All three branches together work best, and every single-key variant lands between the first two rows. Dropping one costs 1.5–3.1 cm, depth most of all. Swapping the future targets for current-frame reconstruction gives back nearly the whole gain — the foresight matters more than the distillation — and a horizon of roughly one action chunk beats both a shorter and a longer one. These are a few centimetres apart, so read them as tendencies.

That ranking is open-loop only. On the physical robot it reverses — the arm without world priors is the strongest of all (Table 5).

Table 4. World-model priors on the approaching set. Each entry is the mean over several checkpoints of one pretraining run.
Armdₕ₁₀ (cm) ↓Prog. (%) ↑θalign (°) ↓
no prior27.640.934.8
three priors22.653.426.5
without the semantics prior24.150.229.2
without the depth prior25.746.131.1
without the dynamics prior24.151.126.8
targets: current frame27.342.733.3
horizon 0.2 s25.845.730.5
horizon 0.8 s24.849.328.3
dynamics target: 3D motion flow25.247.128.6
Test on real robot IRON-R01

On IRON-R01, at 10 Hz, strictly out of distribution

Every real-robot evaluation uses unseen object instances, unseen affordances, or unseen reasoning prompts — ten trials per task, scored for binary success and milestone progress.

Semantic — what to interact with

Unseen categories under heavy clutter: an orange, a toy car, into a basket.

Affordance — where to interact

Grasps dictated by structure: a kettle by the handle, a basket by its handle and handed over.

Reasoning — which one

Prompts resolved before acting: the grapes into the left basket, the blue cup not the yellow.

The rise is strongest in the reasoning and semantic categories; affordance gains less. The scaling runs deliberately leave the world-prior branches off, so that what the curves show is the effect of data scale alone.

Success and progress rate against pretraining hours for six out-of-distribution tasks in three categories.
Figure 10A clear scaling effect on real-robot OOD success and progress. Performance stays flat until 10k pretraining hours, where it jumps: average success is below 12% for every budget up to 5,000 hours and rises above 40% (40–90% per task) at 10,000 hours.

Three ablations of IronMind run against public InternVLA-A1.5 weights post-trained on identical data and compute. Camera space without world priors wins on both metrics — though IronMind and the baseline differ in more than initialisation: IronMind shares one action space across pretraining and post-training, the baseline’s is re-initialised, and the corpora differ.

Table 5. IronMind ablations versus InternVLA-A1.5. Real-robot OOD results on IRON-R01, averaged over the six prompts (10 trials each); same post-training data and compute budget.
InitializationSuccess ↑Progress ↑
InternVLA-A1.5 (public weights)30.0%47.8%
IronMind (ours, torso-space)26.7%46.2%
IronMind (ours, camera-space w/ world prior)31.7%53.3%
IronMind (ours, camera-space w/o world prior)55.0%71.3%

Two lessons from the ablations

Camera space generalises better

Two IronMind models, one in torso reference space and one in camera space, with data scale aligned at 10,000 hours and the same reference frame used in both pretraining and post-training so the comparison is fair. Camera space wins clearly — to our knowledge the first result showing better generalisation from a camera reference frame in ego-centric pretraining for humanoid manipulation.

World priors may not help on the robot

Dropping world-prior supervision improves real-robot performance, even though the open-loop evaluation shows those same branches clearly helping during pretraining (Table 4). The gap between the two is an open question worth chasing.

Paper

Read the paper

IronMind: Scaling Humanoid Dexterous Manipulation via Camera-Space Ego-Centric Pretraining.

Credits

Contributors

XPENG Robotics.

Authors Huimin Pan1*, Yufan Ren1*†, Kunpeng Song1*, Siyang Wang1*, Xiwen Zhang1*, Xiaoyun Hu1,2§, Zhuoxu Duan1,3§, Hanrui Zheng1,4§, Jialeng Ni1,5§, Nathan Zhao1, Sibo Ma1,6§, Zhenxuan Fan1,7§, Zhongyang Che1, Danny Bao1, Jiacheng Wei1, Jerry Bai1, Xiaoyu Yue1, Xiaoyang Guo1, Chenyi Chen1‡ * Equal contribution, listed alphabetically by last name  ·  † Project lead  ·  ‡ Supervision  ·  § Work done during an internship at XPENG Robotics
1 XPENG Robotics  ·  2 Tongji University  ·  3 Rensselaer Polytechnic Institute  ·  4 University of Science and Technology of China  ·  5 University of Michigan  ·  6 University of British Columbia  ·  7 Zhejiang University
Acknowledgements

We thank members of the XPENG Robotics team — Zike Cheng, Ruixuan Zhao, Mengen Luo, Luoyu Bai, Ying Zhang, Haoyue Liu, Peipeng Chen, Ruyi Li, Yizhe Liu, Yongqi Meng, Yinggan Xu, Xi Chen, Bin Shen, and Yuying Ge — for insightful discussions and deployment support, and Zhenglong Du, Hongmei Gan, and Weiqi Liang for assistance with teleoperation and data collection.

@article{ironmind2026,
  title   = {IronMind: Scaling Humanoid Dexterous Manipulation via Camera-Space Ego-Centric Pretraining},
  author  = {Pan, Huimin and Ren, Yufan and Song, Kunpeng and Wang, Siyang and
             Zhang, Xiwen and Hu, Xiaoyun and Duan, Zhuoxu and Zheng, Hanrui and
             Ni, Jialeng and Zhao, Nathan and Ma, Sibo and Fan, Zhenxuan and
             Che, Zhongyang and Bao, Danny and Wei, Jiacheng and Bai, Jerry and
             Yue, Xiaoyu and Guo, Xiaoyang and Chen, Chenyi},
  journal = {XPENG Technical Report},
  year    = {2026}
}