Command Palette
Search for a command to run...
Modélisation d’actions du monde avec planification visuelle progressive
Modélisation d’actions du monde avec planification visuelle progressive
Fei Zhang Zhaochong An Duncan Frost Yikai Wang Pengfei Liu Ya Zhang Michal Drozdzal Amir Bar
Résumé
Les modèles d’action du monde (WAM) se sont imposés comme un paradigme prometteur pour la commande robotique en prédisant conjointement les dynamiques visuelles futures et les actions à partir d’une observation initiale et d’une instruction. Cependant, les WAM existants peinent à réaliser des prédictions à long horizon, car la génération de déploiements vidéo denses est très inefficace. Certains WAM récents remédient à ce problème en prédisant une seule image future sans générer la vidéo complète, mais cette approche néglige la manière de progresser vers l’objectif. Nous présentons ProWAM, un modèle d’action du monde progressif qui prédit conjointement les actions et une séquence ordonnée de sous-objectifs visuels épars, offrant ainsi un guidage visuel explicite pour ancrer la génération d’actions tout au long de l’exécution de la tâche. Cette conception s’étend naturellement, car la prédiction de sous-objectifs peut être apprise à partir de vidéos à grande échelle sans actions, ce qui permet au backbone vidéo de décharger la planification visuelle complexe de la politique d’action. Pour une génération d’actions efficace, ProWAM exécute une seule passe avant du backbone vidéo afin de mettre en cache des caractéristiques éparses de sous-objectifs, éliminant ainsi la génération itérative de la vidéo complète et ne nécessitant qu’un débruitage léger des actions lors de la replanification. À travers de nombreuses évaluations, ProWAM atteint une robustesse supérieure hors distribution. Sur les bancs d’essai en simulation, il établit de nouveaux résultats de pointe sur LIBERO-Plus (85,8 %) et RoboTwin randomisé (75,7 %), dépassant le meilleur modèle de référence avec des gains relatifs allant jusqu’à +35,9 %. Sur RoboCasa365, ProWAM atteint un taux de réussite de 48,1 % et 18,2 % sur le sous-ensemble difficile Composite-Unseen, se classant 4e au classement général. Fait crucial, lors d’expériences réelles en zero-shot, ProWAM atteint 70,0 % de réussite, surpassant le meilleur modèle de référence de +15,0 points (de 55,0 % à 70,0 %, soit un gain relatif de +27,3 %) dans des scènes inédites. Ces résultats démontrent la valeur d’une anticipation visuelle indexée sur la progression pour la commande en boucle fermée.
One-sentence Summary
Researchers from Shanghai Jiao Tong University, SII, Meta, and Imperial College London present ProWAM, a progressive world action model that jointly predicts actions and sparse visual sub-goals to provide explicit visual guidance for robotic control, using a single video-backbone forward pass to cache sub-goal features and lightweight action denoising during replanning; across evaluations it achieves state-of-the-art results on LIBERO-Plus (85.8%) and randomized RoboTwin (75.7%), and attains 70.0% zero-shot real-world success.
Key Contributions
- ProWAM is a progressive world action model that jointly predicts actions and an ordered sequence of sparse visual sub-goals indexed by relative progress, integrating visual planning and physical control within a unified generative framework without a separate high-level planner.
- ProWAM couples cached sub-goal features to action generation through shared attention and uses a single video-backbone forward pass to cache those features, eliminating iterative full-video rollouts and requiring only lightweight action denoising during replanning; the architecture also supports pretraining on large-scale action-free videos followed by joint vision-action fine-tuning.
- Across evaluations, ProWAM sets new state-of-the-art results on LIBERO-Plus (85.8%) and randomized RoboTwin (75.7%), achieves 48.1% success on RoboCasa365 and 18.2% on Composite-Unseen, and reaches 70.0% zero-shot real-world success, outperforming the strongest baseline by 15.0 points or 27.3% relative.
Introduction
Vision-language-action models inherit useful semantic knowledge but remain data-hungry because they depend on large, costly robot demonstration datasets. World action models offer a complementary source of transferable visual-dynamics priors from large-scale video generation, yet prior designs face a clear tradeoff: dense future rollouts provide short-horizon guidance at high computational cost, while efficient zero-imagination or endpoint-only methods lack explicit intermediate-progress guidance over longer tasks. The authors introduce ProWAM, a progressive world action model that represents long-horizon futures as an ordered sequence of sparse visual sub-goals indexed by normalized task progress. ProWAM integrates these progress-conditioned sub-goals into a video generation expert and couples them to a lightweight diffusion-based action expert, enabling efficient closed-loop control without requiring dense video rollouts or an external planner.
Method
To enable explicit visual planning in World Action Models (WAM), the authors introduce the Progressive World Action Model (ProWAM), which generalizes final-goal prediction into a progressive sequence of sparse sub-goals. Unlike classical hierarchical planners that rely on dedicated multi-stage goal-generation modules, ProWAM generates the entire sub-goal trajectory natively and injects these features directly into the action expert.
As shown in the figure below:
The authors parameterize intermediate sub-goals using a normalized progress coordinate r∈[0,1], representing the relative temporal position along the remaining demonstration trajectory from the current observation to its endpoint. Given k monotonically increasing progress levels 0<r1<⋯<rk≤1, the i-th sub-goal corresponds to the frame at:
τ(ri)=τ0+⌊ri(τend−τ0)⌉,i=1,…,k,where τ0 and τend denote the current-frame and demonstration-endpoint indices, respectively. Each selected frame is independently encoded, and the resulting sequence spans selected future states from near-term to longer-term goals. During training, the number of sub-goals k is sampled from a uniform distribution to enable a flexible sub-goal learning strategy.
To incorporate this guidance directly into the architecture, the progressive sub-goal sequence is appended to the video latent sequence:
Z=current observationX0clip transitionsX1,…,Xf/4−1progressive sub-goalsG1,…,Gk,where X denotes the video latent frames. During training, X0 serves as a clean observation latent, while the transition clip and sub-goal latents follow the flow-matching process.
To specify the target temporal stage of each sub-goal, the authors implement progress-conditioned modulation. The progress coordinate ri is mapped through a sinusoidal embedding followed by a multi-layer perceptron to yield an adaptive LayerNorm modulation offset Δmod(ri). This offset is added directly to the standard flow-matching timestep modulation:
mod(Gi)=mod(t)+Δmod(ri).This anchors each sub-goal to its corresponding task stage relative to the reference observation.
For the network architecture, ProWAM implements the video expert and action expert via a mixture-of-transformers architecture with shared attention heads. Information flow is governed by a structured visibility mask. Sub-goal slots attend bidirectionally among themselves and to the reference observation, allowing intermediate milestones to share visual features. Transition latents attend bidirectionally among themselves while conditioning on the reference observation and sub-goals to maintain temporal consistency. Crucially, the video expert does not attend to action tokens, maintaining an action-independent visual forward pathway. Conversely, action tokens attend bidirectionally among themselves and cross-attend to the visual anchors, grounding low-level action execution directly in the predicted sub-goals.
The joint policy is explicitly conditioned on the visual context chain anchored by the current observation. To jointly enable progressive visual imagination and downstream action control, ProWAM is trained via flow matching with independent timesteps for the video and action experts. The full objective combines three velocity-regression terms:
Lfinal=λvideoLvideo+λsubLsub+λactLact,where the first two terms supervise flow-matching velocity predictions for transition and sub-goal latents, and the third trains the action expert to predict velocities for action chunks.
The authors optimize this objective across two distinct phases. In the first stage, they pretrain the video expert on action-free video corpora to learn visual dynamics and progress-conditioned sub-goal prediction. In the second stage, they jointly fine-tune both experts on robot demonstrations, grounding action execution directly in the planned visual sub-goals. This allows ProWAM to leverage vast action-free video corpora for scalable pre-training, improving sample efficiency when fine-tuning on limited robot demonstrations.
For efficient inference, the authors employ sub-goal caching rather than performing burdensome joint video-action denoising. At each replanning step, the video expert is evaluated once to predict the sub-goals, and their key-value features are cached. The action head then reuses this fixed cache throughout the action-denoising steps. Upon receiving a new observation, ProWAM recomputes the sub-goals and refreshes the cache, significantly reducing the number of video-expert forward passes.
Experiment
The experiments evaluate ProWAM across LIBERO, LIBERO-Plus, RoboTwin, RoboCasa365, and zero-shot real-robot tasks, comparing a base variant with a video-pretrained variant. Simulated results show near-saturated in-domain performance and stronger zero-shot generalization, with video sub-goal pretraining primarily improving robustness to visual domain shifts and long-horizon unseen compositions. Real-world zero-shot deployment confirms direct generalization without environment-specific fine-tuning, and failure analysis indicates that task-relevant geometry in generated sub-goals is critical for successful execution. Ablations on RoboTwin demonstrate that cached sub-goal features reduce computation and latency, nearer sub-goal horizons improve out-of-domain robustness, and generated sub-goals preserve useful scene structure despite minor visual artifacts.
On the LIBERO benchmark suite, in-domain success rates are near saturated for the proposed methods and the strongest prior models, so the zero-shot LIBERO-Plus setting provides the main discriminative signal. ProWAM attains the highest zero-shot success rate among compared approaches, outperforming prior VLA and world action model baselines. The variant without sub-goal pretraining shows a clear zero-shot drop despite comparable in-domain performance, indicating that sub-goal pretraining aids out-of-distribution generalization. In-domain LIBERO performance is near saturation for both proposed variants and strong baselines, making LIBERO-Plus a more informative evaluation. ProWAM reaches a new state-of-the-art zero-shot LIBERO-Plus success rate of 85.8%, above previous VLA and world action model baselines. ProWAM-Base falls 5.5 points below ProWAM on zero-shot tasks, linking sub-goal pretraining to stronger out-of-distribution robustness.
On RoboTwin, ProWAM achieves the strongest reported clean and randomized success, with large gains over previous vision-language-action and world-action baselines under visual domain randomization. Prior methods degrade sharply under randomization, notably Fast-WAM, which drops from high clean-scene success to single-digit performance. Ablations show video pretraining and near-term sub-goal planning improve out-of-distribution robustness while clean-scene performance remains stable. ProWAM leads in randomized evaluation and outperforms the strongest prior world-action model by a wide margin. Video pretraining mainly improves robustness to visual domain shifts, while clean-scene success remains comparable.
ProWAM reaches 48.1% overall success on RoboCasa365, placing ahead of ABot-M0.6 but behind Paimon-0, Xiaomi-Robotics-1, and Phasor-m7. Its atomic seen performance is only slightly below the leaders, whereas composite tasks, particularly unseen compositions, show a wider gap. This indicates stronger relative performance on shorter-horizon tasks and more difficulty on long-horizon generalization. ProWAM is close to the best methods on atomic seen tasks, trailing the leaders by only a few percentage points. On composite unseen tasks, ProWAM outperforms ABot-M0.6 but remains far behind the top three public methods, with the largest gap to Paimon-0.
In real-world zero-shot evaluation across four physical tasks, ProWAM achieves the highest average success rate among compared methods, outperforming the larger DreamZero model by a clear margin. Its advantage is especially strong on spatially grounded placement, where it reaches perfect success, demonstrating direct generalization without environment-specific fine-tuning. ProWAM uses fewer parameters than DreamZero yet achieves a higher average success rate across the four real-world tasks. ProWAM reaches perfect success on spatially grounded placement, outperforming the baselines on that task. The real-world tasks span pick-and-place, long-horizon multi-object assembly, contact-rich placement, and spatially grounded placement.
Across four evaluation settings, ProWAM is tested on zero-shot LIBERO-Plus generalization, RoboTwin visual domain randomization, RoboCasa365 multi-task performance, and real-world zero-shot manipulation. It achieves state-of-the-art zero-shot LIBERO-Plus success and strong randomized RoboTwin results, with ablations showing that sub-goal pretraining and video pretraining improve out-of-distribution robustness. On RoboCasa365 it performs near the leaders on atomic seen tasks but lags on composite unseen tasks, indicating stronger short-horizon ability. In real-world evaluation it attains the highest average success with fewer parameters than baselines and reaches perfect success on spatially grounded placement.