WholeBodyWAMGeneralizing Pre-trained World-Action Priors to Humanoid Loco-Manipulation via WBC-Grounded Coordination

Zhuo Li1,4,† Yiming Yao2,4,* Jim Tan3,4,* Mengjie Jing1,4 Zhipeng Dong1,4 Fei Chen1,4,‡
1The Chinese University of Hong Kong 2The University of Hong Kong 3Peking University 4Φ-Institute

† Project Lead‡ Corresponding Author* Equal Contribution

From manipulation priors to coordinated whole-body behavior. WholeBodyWAM enables a Unitree G1 humanoid to coordinate walking, reaching, and manipulation. CartServe demonstration, shown at 2× playback.

Abstract

World Action Models (WAMs) offer a promising approach to general-purpose robot manipulation by jointly modeling visual dynamics and actions. However, most WAM studies focus on tabletop or arm-centric manipulation, while humanoid loco-manipulation remains less explored. To address this gap, we introduce WholeBodyWAM, which jointly predicts future visual dynamics, manipulation actions, and whole-body control intents for generalizable humanoid loco-manipulation. It preserves pre-trained world-action priors while grounding heterogeneous whole-body controller (WBC) semantics and coordinating whole-body behavior. Extensive experiments show that WholeBodyWAM achieves an overall simulation task success rate of 91.9%, with a 0.23 improvement in real-world out-of-distribution task progress and a 70% reduction in success-rate variance across WBCs relative to the respective baselines. These results suggest a path toward scalable humanoid whole-body intelligence by extending pre-trained world-action priors through structured WBC grounding and coordination, rather than relearning whole-body behavior from scratch.

91.9%Simulation task success
+0.23Real-world OOD task progress
over DreamZero
70%Lower cross-WBC variance
than Cosmos-3

Overview

WholeBodyWAM preserves world-action priors and jointly predicts visual dynamics, manipulation actions, and whole-body controller commands, evaluated on six simulation and eight real-world tasks.
Figure 1. Overview of WholeBodyWAM. Given language, visual history, and proprioception, the model jointly predicts future visual dynamics, manipulation intent, and unified WBC commands. Pre-trained world-action priors are preserved and generalized through structured WBC grounding and coordination.

Method

WholeBodyWAM extends a pre-trained world action model with a structured whole-body control stream. Visual dynamics, manipulation actions, and controller-facing intent are generated jointly in one shared Diffusion Transformer.

Model architecture with visual, manipulation, and UWBC token streams, plus a manipulability-informed gate that controls coordination-aware self-attention.
Figure 2. Model architecture. A shared Diffusion Transformer processes three typed streams. Coordination-Aware Self-Attention (CASA) strengthens manipulation-to-UWBC attention when task-directional arm manipulability decreases.
01

Preserve manipulation priors

Structured action factorization retains the pre-trained visual–manipulation pathway and introduces distinct UWBC tokens, making whole-body intent explicitly addressable.

02

Ground controller semantics

The Unified WBC Interface organizes 56 command slots: 46 shared slots and 10 controller-specific residual slots. Unsupported or undefined fields are masked according to controller capabilities.

03

Coordinate the whole body

CASA uses a manipulability-informed gate to strengthen manipulation-to-UWBC information flow, allowing manipulation intent to guide compensatory body motion.

WholeBodyWAM predicts task-level intent; the downstream WBC handles balance, gait, contact, and low-level tracking. Evaluated controllers: SONIC, AMO, and GEAR WBC.

Real-World Demonstrations

Eight tasks on a Unitree G1 humanoid span manipulation-dominant and coordination-intensive behaviors, from pouring and watering to carrying objects and opening doors.

CartServe

Push the cart and serve a snack at the side table.

2× playback

BoxTransfer

Crouch, lift a box, and carry it to the target location.

2× playback

BasketCarry

Lift the basket, turn, carry it, and lower it into place.

2× playback

DoorEntry

Open the door and walk into the room.

2× playback

TowelPlace

Pick up the towel and place it in the bin.

2× playback

TableCleanup

Lift the container, wipe the table, and discard the cloth.

2× playback

PlantWater

Water the plant and return the watering can.

2× playback

TeapotPour

Grasp the teapot, pour into the selected cup, and return it.

Source timing preserved

Generalization Beyond Training

TowelPlace · Changing scene layout

With the towel support displaced 10 cm farther away, the first grasp fails. The robot steps forward and adjusts its body to successfully grasp the towel on the next attempt.

1× source playback · Sequence associated with Figure 8

TeapotPour · Changing human preference

When the participant indicates a different cup during execution, the robot redirects the teapot toward the newly selected cup, then returns and releases the teapot.

Source timing preserved

Experimental Results

Evaluation covers six simulation tasks, eight real-world tasks, and three WBC interfaces. Results measure task success, normalized task progress, and sensitivity to the downstream controller.

Simulation performance

With SONIC in SIMPLE, WholeBodyWAM achieves 91.9% overall task success across six tasks and three perturbation levels. Removing any of its three main components lowers overall success.

Table II. Simulation task success rate (%). Each task reports L0 / L1 / L2.
MethodMovePickBendPickHandoverPickBetweenTablesTabletopGraspMoveBendPickOverall
DreamZero70 / 65 / 5575 / 70 / 6085 / 80 / 7050 / 45 / 3590 / 85 / 7560 / 55 / 4565.0
DreamZero-PT80 / 75 / 6585 / 80 / 7090 / 85 / 8060 / 55 / 4595 / 90 / 8070 / 65 / 5573.6
Cosmos-390 / 85 / 8095 / 90 / 80100 / 95 / 9085 / 80 / 70100 / 95 / 8585 / 80 / 7086.4
GR00T N1.655 / 50 / 4065 / 60 / 5060 / 55 / 4530 / 25 / 1575 / 70 / 6040 / 35 / 2547.5
Ψ₀90 / 85 / 7590 / 85 / 8095 / 90 / 8075 / 70 / 60100 / 95 / 9080 / 75 / 6582.2
WholeBodyWAM95 / 90 / 85100 / 95 / 90100 / 95 / 9095 / 90 / 80100 / 95 / 8595 / 90 / 8591.9
w/o CASA90 / 85 / 8095 / 90 / 85100 / 95 / 8585 / 80 / 7595 / 90 / 8590 / 85 / 7586.9
w/o UWBC90 / 85 / 7595 / 90 / 8095 / 90 / 8585 / 80 / 7095 / 90 / 8585 / 80 / 7084.7
w/o SAF85 / 80 / 7090 / 85 / 7595 / 90 / 8080 / 75 / 6595 / 90 / 8080 / 75 / 6580.8

100 task-specific fine-tuning demonstrations per task. SAF: structured action factorization; UWBC: Unified WBC Interface; CASA: Coordination-Aware Self-Attention.

Real-world performance

WholeBodyWAM achieves 81.3% in-distribution (ID) and 68.8% out-of-distribution (OOD) success, compared with 57.5% and 40.0% for DreamZero. Mean OOD task progress improves from 0.59 to 0.82.

Table III. Real-world task success rate (%) over 20 trials per method and task.
TaskWholeBodyWAM
ID
DreamZero
ID
WholeBodyWAM
OOD
DreamZero
OOD
TowelPlace85657550
BasketCarry75405515
TeapotPour85758065
BoxTransfer75356020
CartServe80506530
DoorEntry80457025
TableCleanup90858070
PlantWater80656545
Overall81.357.568.840.0

OOD evaluation changes object configurations and language instructions. Overall values are reported as in the manuscript.

Cross-WBC results: WholeBodyWAM achieves 89.2 percent mean task success and 10.5 squared percentage-point variance, versus 80.2 percent and 35.0 for Cosmos-3.
Figure 4. Cross-WBC robustness. WholeBodyWAM achieves the highest mean task success and lowest variance across SONIC, AMO, and GEAR WBC. Each controller receives separate task-specific fine-tuning.
Task success across 50, 100, 200, and 300 fine-tuning demonstrations per task, showing WholeBodyWAM's advantage with fewer demonstrations.
Figure 5. Data efficiency. The largest advantage appears with few task-specific demonstrations. These fine-tuning budgets are separate from the 15K whole-body post-training demonstrations.
Radar plots comparing normalized task progress for WholeBodyWAM and DreamZero across eight real-world tasks under ID and OOD conditions.
Figure 7. Task progress under ID and OOD conditions. Normalized progress measures how far execution proceeds through the ordered task stages before the first failure or irreversible deviation.

Read the Paper

Read the complete method, evaluation protocol, ablations, and references in the eight-page manuscript.

BibTeX

@misc{li2026wholebodywam,
  title={{WholeBodyWAM}: Generalizing Pre-trained World-Action Priors to
         Humanoid Loco-Manipulation via {WBC}-Grounded Coordination},
  author={Zhuo Li and Yiming Yao and Jim Tan and Mengjie Jing and
          Zhipeng Dong and Fei Chen},
  year={2026},
  eprint={2609.16644},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  doi={10.48550/arXiv.2609.16644},
  url={https://arxiv.org/abs/2609.16644}
}