Command Palette
Search for a command to run...
漸進的視覚プランニングによる世界行動モデリング
漸進的視覚プランニングによる世界行動モデリング
Fei Zhang Zhaochong An Duncan Frost Yikai Wang Pengfei Liu Ya Zhang Michal Drozdzal Amir Bar
概要
世界行動モデル(WAM)は、初期観察と指示から将来の視覚的ダイナミクスと行動を同時に予測することにより、ロボット制御の有望なパラダイムとして台頭している。しかし、既存のWAMは、密なビデオロールアウトの生成が非常に非効率であるため、長期間の予測に苦戦している。最近の一部のWAMは、全ビデオを生成せずに単一の将来フレームを予測することでこれに対処しているが、このアプローチは目標に向かってどのように進行するかを無視している。我々は、行動と順序付けられた疎な視覚的サブゴール系列を同時に予測する漸進的世界行動モデルProWAMを提案する。これにより、タスク実行全体を通じて行動生成を固定する明示的な視覚的ガイダンスを提供する。この設計は自然にスケールし、サブゴール予測は大規模な行動ラベルなし動画から学習できるため、ビデオバックボーンが複雑な視覚的プランニングを行動方策からオフロードできる。効率的な行動生成のため、ProWAMは1回のビデオバックボーン順伝播を実行して疎なサブゴール特徴をキャッシュし、反復的な全ビデオ生成を排除して、再計画時に軽量な行動デノイジングのみを必要とする。広範な評価を通じて、ProWAMは優れた分布外ロバスト性を達成する。シミュレーションベンチマークでは、LIBERO-Plus(85.8%)とランダム化RoboTwin(75.7%)で新たな最先端結果を樹立し、最強のベースラインを最大+35.9%の相対的向上で上回った。RoboCasa365では、ProWAMは48.1%の成功率を達成し、困難なComposite-Unseen分割で18.2%を達成して総合4位に入った。重要なことに、ゼロショット実世界実験では、ProWAMは新しい場面で70.0%の成功率を達成し、最強のベースラインを+15.0ポイント上回った(55.0%から70.0%、相対的向上+27.3%)。これらの結果は、閉ループ制御における進捗指標付き視覚的先見の価値を実証している。
One-sentence Summary
Researchers from Shanghai Jiao Tong University, SII, Meta, and Imperial College London present ProWAM, a progressive world action model that jointly predicts actions and sparse visual sub-goals to provide explicit visual guidance for robotic control, using a single video-backbone forward pass to cache sub-goal features and lightweight action denoising during replanning; across evaluations it achieves state-of-the-art results on LIBERO-Plus (85.8%) and randomized RoboTwin (75.7%), and attains 70.0% zero-shot real-world success.
Key Contributions
- ProWAM is a progressive world action model that jointly predicts actions and an ordered sequence of sparse visual sub-goals indexed by relative progress, integrating visual planning and physical control within a unified generative framework without a separate high-level planner.
- ProWAM couples cached sub-goal features to action generation through shared attention and uses a single video-backbone forward pass to cache those features, eliminating iterative full-video rollouts and requiring only lightweight action denoising during replanning; the architecture also supports pretraining on large-scale action-free videos followed by joint vision-action fine-tuning.
- Across evaluations, ProWAM sets new state-of-the-art results on LIBERO-Plus (85.8%) and randomized RoboTwin (75.7%), achieves 48.1% success on RoboCasa365 and 18.2% on Composite-Unseen, and reaches 70.0% zero-shot real-world success, outperforming the strongest baseline by 15.0 points or 27.3% relative.
Introduction
Vision-language-action models inherit useful semantic knowledge but remain data-hungry because they depend on large, costly robot demonstration datasets. World action models offer a complementary source of transferable visual-dynamics priors from large-scale video generation, yet prior designs face a clear tradeoff: dense future rollouts provide short-horizon guidance at high computational cost, while efficient zero-imagination or endpoint-only methods lack explicit intermediate-progress guidance over longer tasks. The authors introduce ProWAM, a progressive world action model that represents long-horizon futures as an ordered sequence of sparse visual sub-goals indexed by normalized task progress. ProWAM integrates these progress-conditioned sub-goals into a video generation expert and couples them to a lightweight diffusion-based action expert, enabling efficient closed-loop control without requiring dense video rollouts or an external planner.
Method
To enable explicit visual planning in World Action Models (WAM), the authors introduce the Progressive World Action Model (ProWAM), which generalizes final-goal prediction into a progressive sequence of sparse sub-goals. Unlike classical hierarchical planners that rely on dedicated multi-stage goal-generation modules, ProWAM generates the entire sub-goal trajectory natively and injects these features directly into the action expert.
As shown in the figure below:
The authors parameterize intermediate sub-goals using a normalized progress coordinate r∈[0,1], representing the relative temporal position along the remaining demonstration trajectory from the current observation to its endpoint. Given k monotonically increasing progress levels 0<r1<⋯<rk≤1, the i-th sub-goal corresponds to the frame at:
τ(ri)=τ0+⌊ri(τend−τ0)⌉,i=1,…,k,where τ0 and τend denote the current-frame and demonstration-endpoint indices, respectively. Each selected frame is independently encoded, and the resulting sequence spans selected future states from near-term to longer-term goals. During training, the number of sub-goals k is sampled from a uniform distribution to enable a flexible sub-goal learning strategy.
To incorporate this guidance directly into the architecture, the progressive sub-goal sequence is appended to the video latent sequence:
Z=current observationX0clip transitionsX1,…,Xf/4−1progressive sub-goalsG1,…,Gk,where X denotes the video latent frames. During training, X0 serves as a clean observation latent, while the transition clip and sub-goal latents follow the flow-matching process.
To specify the target temporal stage of each sub-goal, the authors implement progress-conditioned modulation. The progress coordinate ri is mapped through a sinusoidal embedding followed by a multi-layer perceptron to yield an adaptive LayerNorm modulation offset Δmod(ri). This offset is added directly to the standard flow-matching timestep modulation:
mod(Gi)=mod(t)+Δmod(ri).This anchors each sub-goal to its corresponding task stage relative to the reference observation.
For the network architecture, ProWAM implements the video expert and action expert via a mixture-of-transformers architecture with shared attention heads. Information flow is governed by a structured visibility mask. Sub-goal slots attend bidirectionally among themselves and to the reference observation, allowing intermediate milestones to share visual features. Transition latents attend bidirectionally among themselves while conditioning on the reference observation and sub-goals to maintain temporal consistency. Crucially, the video expert does not attend to action tokens, maintaining an action-independent visual forward pathway. Conversely, action tokens attend bidirectionally among themselves and cross-attend to the visual anchors, grounding low-level action execution directly in the predicted sub-goals.
The joint policy is explicitly conditioned on the visual context chain anchored by the current observation. To jointly enable progressive visual imagination and downstream action control, ProWAM is trained via flow matching with independent timesteps for the video and action experts. The full objective combines three velocity-regression terms:
Lfinal=λvideoLvideo+λsubLsub+λactLact,where the first two terms supervise flow-matching velocity predictions for transition and sub-goal latents, and the third trains the action expert to predict velocities for action chunks.
The authors optimize this objective across two distinct phases. In the first stage, they pretrain the video expert on action-free video corpora to learn visual dynamics and progress-conditioned sub-goal prediction. In the second stage, they jointly fine-tune both experts on robot demonstrations, grounding action execution directly in the planned visual sub-goals. This allows ProWAM to leverage vast action-free video corpora for scalable pre-training, improving sample efficiency when fine-tuning on limited robot demonstrations.
For efficient inference, the authors employ sub-goal caching rather than performing burdensome joint video-action denoising. At each replanning step, the video expert is evaluated once to predict the sub-goals, and their key-value features are cached. The action head then reuses this fixed cache throughout the action-denoising steps. Upon receiving a new observation, ProWAM recomputes the sub-goals and refreshes the cache, significantly reducing the number of video-expert forward passes.
Experiment
The experiments evaluate ProWAM across LIBERO, LIBERO-Plus, RoboTwin, RoboCasa365, and zero-shot real-robot tasks, comparing a base variant with a video-pretrained variant. Simulated results show near-saturated in-domain performance and stronger zero-shot generalization, with video sub-goal pretraining primarily improving robustness to visual domain shifts and long-horizon unseen compositions. Real-world zero-shot deployment confirms direct generalization without environment-specific fine-tuning, and failure analysis indicates that task-relevant geometry in generated sub-goals is critical for successful execution. Ablations on RoboTwin demonstrate that cached sub-goal features reduce computation and latency, nearer sub-goal horizons improve out-of-domain robustness, and generated sub-goals preserve useful scene structure despite minor visual artifacts.
On the LIBERO benchmark suite, in-domain success rates are near saturated for the proposed methods and the strongest prior models, so the zero-shot LIBERO-Plus setting provides the main discriminative signal. ProWAM attains the highest zero-shot success rate among compared approaches, outperforming prior VLA and world action model baselines. The variant without sub-goal pretraining shows a clear zero-shot drop despite comparable in-domain performance, indicating that sub-goal pretraining aids out-of-distribution generalization. In-domain LIBERO performance is near saturation for both proposed variants and strong baselines, making LIBERO-Plus a more informative evaluation. ProWAM reaches a new state-of-the-art zero-shot LIBERO-Plus success rate of 85.8%, above previous VLA and world action model baselines. ProWAM-Base falls 5.5 points below ProWAM on zero-shot tasks, linking sub-goal pretraining to stronger out-of-distribution robustness.
On RoboTwin, ProWAM achieves the strongest reported clean and randomized success, with large gains over previous vision-language-action and world-action baselines under visual domain randomization. Prior methods degrade sharply under randomization, notably Fast-WAM, which drops from high clean-scene success to single-digit performance. Ablations show video pretraining and near-term sub-goal planning improve out-of-distribution robustness while clean-scene performance remains stable. ProWAM leads in randomized evaluation and outperforms the strongest prior world-action model by a wide margin. Video pretraining mainly improves robustness to visual domain shifts, while clean-scene success remains comparable.
ProWAM reaches 48.1% overall success on RoboCasa365, placing ahead of ABot-M0.6 but behind Paimon-0, Xiaomi-Robotics-1, and Phasor-m7. Its atomic seen performance is only slightly below the leaders, whereas composite tasks, particularly unseen compositions, show a wider gap. This indicates stronger relative performance on shorter-horizon tasks and more difficulty on long-horizon generalization. ProWAM is close to the best methods on atomic seen tasks, trailing the leaders by only a few percentage points. On composite unseen tasks, ProWAM outperforms ABot-M0.6 but remains far behind the top three public methods, with the largest gap to Paimon-0.
In real-world zero-shot evaluation across four physical tasks, ProWAM achieves the highest average success rate among compared methods, outperforming the larger DreamZero model by a clear margin. Its advantage is especially strong on spatially grounded placement, where it reaches perfect success, demonstrating direct generalization without environment-specific fine-tuning. ProWAM uses fewer parameters than DreamZero yet achieves a higher average success rate across the four real-world tasks. ProWAM reaches perfect success on spatially grounded placement, outperforming the baselines on that task. The real-world tasks span pick-and-place, long-horizon multi-object assembly, contact-rich placement, and spatially grounded placement.
Across four evaluation settings, ProWAM is tested on zero-shot LIBERO-Plus generalization, RoboTwin visual domain randomization, RoboCasa365 multi-task performance, and real-world zero-shot manipulation. It achieves state-of-the-art zero-shot LIBERO-Plus success and strong randomized RoboTwin results, with ablations showing that sub-goal pretraining and video pretraining improve out-of-distribution robustness. On RoboCasa365 it performs near the leaders on atomic seen tasks but lags on composite unseen tasks, indicating stronger short-horizon ability. In real-world evaluation it attains the highest average success with fewer parameters than baselines and reaches perfect success on spatially grounded placement.