XPENG Robotics Foundation Model Team

IRON

Built in-house.
Engineered end to end.

Iron is XPENG Robotics Foundation Model Team’s in-house model family—language, vision, and multimodal—built end to end from data to deployment. Explore each model below.

3 model lines
Full-stack in-house R&D
Edge-ready deployment
02 / NEWS

News & updates

Model releases and product announcements—version history for Iron as it ships and iterates.

03 / WHY IRON

Why Iron

Control at every layer.

IronLLM-0.6B is more than a checkpoint. Three coordinated strengths let us improve capability, efficiency, and deployment as one system.

01DATA

In-house pipeline,
unified control.

Training data combines in-house, open-source, and synthetic sources—processed through our pipeline for quality control and mixture so the final corpus stays consistent.

Unifieddata pipeline
02MODEL DESIGN

Edge-native
by design.

Architected natively for on-device scenarios: a hybrid attention backbone for efficient long context, an RMSNorm-free Light variant for extreme deployment, and X-MTP multi-token prediction for faster decoding.

1.48×Peak X-MTP speedup
03FULL STACK

Pre. Mid. Post.
One loop.

Capability is shaped end to end—broad pre-training, long-context mid-training, then general SFT, domain-specialist training, and multi-domain distillation—so the recipe, not only the checkpoint, is owned in-house.

3training phases
04 / MODEL

The model

One model. A complete stack.

IronLLM-0.6B

Compact in scale,
complete in capability.

Designed and trained in-house as a practical foundation for knowledge, reasoning, long context, and on-device workloads.

6.2T pretrain tokens
0.6B parameters
64K context
Model architecture

Hybrid by design.
Practical by intent.

A visualization-first view of the structure—optimized for on-device long-context scenarios with lower memory footprint and faster inference.

IronLLM-0.6B hybrid architecture overview
Figure A · IronLLM-0.6B architecture
01

Hybrid Attention Architecture

Optimized for on-device long-context scenarios, we employ a 3:1 ratio of Gated DeltaNet to standard Attention across all 24 layers. This design retains global receptive fields while drastically reducing KV Cache consumption. It delivers faster on-device inference and a lower memory footprint, enabling our 0.6B model to handle complex long-context tasks smoothly on mobile devices.

02

Ultra-Efficient Light Variant

A streamlined, RMSNorm-free variant built for extreme edge deployment. Dynamic Tanh (DyT) replaces every RMSNorm layer, while upper-bounded ReLUx activations and data-independent gating remove quantization-unfriendly sigmoid paths—maximizing inference throughput and quantization efficiency with minimal capability trade-off.

03

X-MTP & Verification Head

Our lightweight multi-token prediction module drafts multiple future tokens while sharing attention KV across prediction steps, avoiding repeated KV-cache replay. A lightweight verification head estimates confidence to adaptively accept drafts without rollback, accelerating decoding.

05 / TRAIN

Training

One model. Three training phases.

Pre-training builds broad capability, mid-training extends usable context, and post-training develops a general instruction model, strengthens specialist capabilities, and integrates them through on-policy distillation.

PRE-TRAINING

A three-stage recipe that shifts from broad web knowledge to STEM density, then reasoning refinement.

STAGE 01

Broad Knowledge Foundation

Establishes a robust linguistic and world knowledge base using massive general web corpora. This phase utilizes a learning rate warmup to stabilize early training dynamics, ensuring broad coverage of fundamental concepts and language patterns.

data
General Web dominant
tokens
4.2T
lr
warmup → max
STAGE 02

Specialized Capability Injection

Significantly increases the density of high-quality STEM data (Math & Code). By maintaining a stable maximum learning rate, this stage effectively injects complex reasoning capabilities and logical problem-solving skills into the model’s foundation without catastrophic forgetting.

data
High-density Math & Code
tokens
1T
lr
constant max
STAGE 03

Reasoning Refinement

Maximizes the proportion of reasoning-intensive samples to sharpen logical precision. The learning rate decay schedule is applied here to consolidate learned capabilities, smoothing the loss landscape and ensuring stable convergence for superior downstream performance.

data
High quality & reasoning-focused
tokens
1T
lr
decay → zero

We adopt a three-stage training pipeline that progressively shifts from general web data to math, code, and reasoning-focused corpora, following a warmup–stable–decay learning rate schedule. Despite training on 6.2T tokens—less than 20% of Qwen3’s 36T pre-training budget—our base model achieves comparable performance to Qwen3.5-0.8B-Base across various core metrics.

MID-TRAINING

We extend usable context while preserving short-context capability from pretraining, then strengthen specialized long-context skills with targeted synthetic data.

01

Context Extension

Two-stage long-context training extends the window from 4K to 32K then 64K, with a staged RoPE-base increase so longer context lands without erasing pretraining. Stages use about 42B and 21B tokens; stretching past 64K added little in our tests.

staged RoPE · two-stage extension
02

Balanced Mixture

To keep short-context strength from pretraining, we stay close to the final pretraining mixture and only upsample selected buckets for long-context learning.

We split the corpus into length-based buckets and tune sampling ratios with ablations, so long-context gains do not trade away competitive short-context benchmarks.

length buckets · replay mix
03

Long-data Curation

Skills such as retrieval and counting are hard to learn from natural long corpora alone, so we curate synthetic long-context data aimed at those gaps.

Used in mid-training or SFT, this data lifts the target long-context abilities without degrading general or short-context benchmarks.

synthetic long data · no short-context tradeoff
01 / GENERAL SFT

General Instruction Foundation

We fine-tune the base model on several million high-quality, multi-domain conversation samples, including general QA, mathematics, code, and tool use.

This stage establishes broad conversational and instruction-following ability as the foundation for the specialist models and unified model that follow.

multi-domain conversations · assistant-response training
02 / DOMAIN SPECIALISTS

Targeted Expertise

Starting from the General SFT model, we independently train specialists for mathematics, code, instruction following, multiple-choice QA, and open-ended QA.

Each domain uses a tailored recipe of supervised fine-tuning, reinforcement learning, or both. RL uses GRPO with domain-specific verifiable rewards or a preference reward model.

five capability domains · tailored SFT / RL
03 / MOPD

Unifying Specialist Capabilities

Multi-Domain On-Policy Distillation (MOPD) transfers complementary capabilities from the domain specialists back into the General SFT model, producing one unified assistant.

The student learns from specialist guidance on its own generated responses, while domain verifiers provide sequence-level rewards. This integrates specialist strengths while preserving broad general ability.

Light variant: skips specialist training and uses the IronLLM-0.6B specialists for MOPD.

student rollouts · specialist guidance · verifiable rewards
IronLLM post-training pipeline from SFT to MOPD
Figure C · Post-training pipeline
06 / DATA

Data engine

Data is the first model.

Our training corpus is built from a combination of in-house generated data, high-quality open-source datasets, and synthetic data.

All data sources are processed through a unified data pipeline, including quality filtering, deduplication, data enhancement, and mixture optimization, to ensure consistent quality across the final training corpus.

Our in-house pipeline enables continuous acquisition, processing, and optimization of large-scale training data, providing a scalable and reproducible data production capability.

IronLLM unified data pipeline
Figure D · Data pipeline
07 / RESULTS

Evaluation results

Small footprint. Serious capability.

Selected post-training results for IronLLM-0.6B, followed by category-level comparisons and detailed benchmark scores.

IronLLM-0.6B highlights

Selected scores · non-thinking

75.6 IFEval

Instruction following

42.1 MMLU-Pro

General knowledge

78.6 GSM8K

Mathematics

49.4 BFCL v3

Function calling

Model comparison

Category averages · non-thinking · hover a bar for the score.

  • Iron 0.6B
  • Iron-Light 0.6B
  • Qwen3 0.6B
  • LFM2 700M
  • Qwen3.5 0.8B
  • MiniCPM 1B

Benchmark details

Bold marks the best non-thinking score shown. Expanded cells show non-thinking / thinking for Qwen3, Qwen3.5, and MiniCPM5.