Command Palette
Search for a command to run...
InternW0-∆: Ein World-Action-Modell, das prädiktive Dynamik und Aktionen mit über 20.000 Stunden offenen Daten verbindet
InternW0-∆: Ein World-Action-Modell, das prädiktive Dynamik und Aktionen mit über 20.000 Stunden offenen Daten verbindet
Zusammenfassung
World Action Models (WAMs) haben sich als vielversprechendes Paradigma für die generalistische Roboter-Manipulation etabliert, indem sie visuelle Dynamik und Aktionsgenerierung gemeinsam modellieren. Eine zentrale Herausforderung besteht darin, komplementäre Priors aus großen vortrainierten Modellen – darunter visuelle Dynamik, Szenensemantik sowie geometrisches und Bewegungsverständnis – wirksam in ein einheitliches Framework für die Roboter-Aktionsgenerierung zu integrieren. Wir stellen InternW0-∆ vor, ein einheitliches World-Action-Modell, das diese Herausforderung adressiert: Vortrainiert auf einem großen heterogenen Korpus übertrifft es frühere Methoden in verschiedenen Simulationsbenchmarks und auf realen Roboterplattformen. InternW0-∆ führt vortrainierte visuelle Dynamik, szenenbezogenes semantisches Verständnis, geometrische und Bewegungspriors aus 4D sowie die Aktionsgenerierung in einem Mixture-of-Transformers-(MoT-)Framework zusammen. Innerhalb des World–Action-MoT interagieren ein vortrainierter Videoexperte und ein Aktionsexperte unter szenenverankerter semantischer Führung durch ein eingefrorenes VLM, während ein vortrainiertes 4D-Foundation-Modell geometrische und Bewegungspriors durch ausschließlich trainingsseitige Destillation einbringt. Um prädiktive visuelle Dynamik in für die Aktionsvorhersage nützliche Repräsentationen zu übersetzen, führen wir Causal Imprint ein, das zukunftsrelevante Szenenveränderungen aus ausschließlich trainingsseitiger Zukunftssupervision lernt und diese prädiktiven Repräsentationen dem Aktionsexperten direkt verfügbar macht, ohne dass bei der Inferenz ein Future-Video-Rollout erforderlich ist. Zur Unterstützung groß angelegten gemeinsamen Trainings konstruieren wir einen heterogenen Korpus, der Roboter-Demonstrationen, UMI-Daten, egozentrische menschliche Demonstrationen und Ego2Robot-Daten umfasst – sorgfältig kuratiert und gefiltert, unter einer gemeinsamen Zustands-Aktions-Repräsentation vereinheitlicht und zeitlich ausgerichtet – und damit über 20.000 Stunden verarbeitete Trainingsdaten liefert; unseres Wissens der größte offene Korpus dieser Art. Wir trainieren InternW0-∆ auf diesem heterogenen Korpus vor und demonstrieren starke Leistung über verschiedene Simulationsbenchmarks und reale Roboterplattformen hinweg. Wir werden Trainingscode, Modellgewichte, Infrastruktur und Datenverarbeitungspipeline sowie – soweit Lizenzen dies erlauben – verarbeitete Daten als Open Source bereitstellen, um den Fortschritt in der verkörperten Intelligenz und physischen KI zu beschleunigen.
One-sentence Summary
Researchers from the Physical Intelligence Team at Shanghai AI Laboratory introduce InternW0-∆, a unified World Action Model that combines a Mixture-of-Transformers framework and Causal Imprint to integrate pretrained video, action, semantic, and 4D geometric priors, training on a 20K+ hour heterogeneous open corpus and outperforming prior methods across simulation and real-robot benchmarks.
Key Contributions
- The paper introduces InternW0-∆, a unified World Action Model that combines pretrained visual dynamics, scene-level semantic understanding, 4D geometric and motion priors, and action generation within a Mixture-of-Transformers framework, outperforming prior methods across diverse simulation benchmarks and real-robot platforms.
- Causal Imprint learns future-relevant scene changes from training-only future supervision and exposes these predictive representations to the action expert, enabling benefits from predictive visual dynamics without future-video rollout at inference.
- The work constructs a curated, filtered, and temporally aligned pretraining corpus of over 20,000 hours spanning robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data, unified under common state-action representations, and transfers the resulting model across simulation benchmarks, real-robot platforms, gripper-based manipulation, and dexterous-hand settings.
Introduction
World Action Models are a promising direction for generalist robot manipulation because they jointly model visual dynamics and robot actions, and large-scale video pretraining can supply useful visual and temporal priors. However, the ability to predict future observations does not directly translate into effective control: action generation also requires identifying task-relevant changes, reasoning about object geometry and motion, and grounding these cues in the current instruction and scene. Prior WAMs often rely on future video generation during inference or do not give the action path direct access to predictive representations. To address this, the authors introduce InternW0-∆, a directed world-action architecture that couples a pretrained video expert with an action expert, using Causal Imprint and training-only 4D-aware distillation to inject action-relevant future, geometric, and motion information without requiring future-video sampling at inference.
Dataset
Dataset overview
The authors use a multi-source embodied dataset for joint training. It combines robot manipulation demonstrations, egocentric human videos, ego-to-robot converted data, and UMI data. All trajectories are unified under one canonical state-action representation.
Unified representation
- All robot actions are mapped into a shared 80-dimensional action space with fixed semantic slots.
- Robot state keeps absolute quantities: current joint positions, absolute end-effector pose, and current gripper or hand configuration.
- End-effector pose is represented as a 3D position plus a continuous 6D rotation.
- For parallel-jaw grippers, hand state uses gripper opening; for dexterous hands, it uses hand joint configuration.
- Actions distinguish joint-space absolute target joint positions from task-space relative end-effector motion.
- Task-space end-effector actions use a 3D translation increment and a 3D rotation vector in the current end-effector local frame.
- An action segment contains Ha consecutive canonical action vectors.
Robot manipulation data
The authors integrate 15 robot datasets from simulated and real-world environments. Reported source details include:
- AgiBotWorld: over 1 million trajectories, about 3,000 hours, dual-arm humanoid data, 217 tasks, multimodal cameras and tactile sensors.
- InternData-A1: over 630K trajectories, 7,400 hours, four embodiments, 70 tasks, 227 scenes, simulated and real.
- RoboMIND / RoboMIND 2.0: 107K trajectories across 479 tasks; version 2.0 expands to over 310K dual-arm trajectories, six embodiments, 739 tasks.
- RW-RL Dataset: over 1,000 hours of real interaction, combining teleoperated demonstrations, autonomous rollouts, interventions, rewards, and termination labels.
- Dexora: 100K simulated plus 10K real bimanual dexterous manipulation trajectories.
- ABC-130K: over 130K real bimanual episodes, over 3,500 hours, 195 tasks.
- Galaxea Open-World Dataset: over 500 hours of real mobile manipulation with subtask-level language annotations.
- RoboCOIN: over 180K demonstrations across 15 robot platforms, 421 tasks, 16 scenarios.
- RH20T: over 110K contact-rich sequences, 147 tasks, with force, audio, depth, and partially tactile data.
- RDT-1B: aggregates 46 datasets into over 1 million episodes, plus over 6K ALOHA demonstrations.
- RoboSet: about 31K real kitchen trajectories across 40 tasks, plus simulated trajectories.
- RealSource-World: over 14 million frames, over 11K dual-arm episodes, 35 tasks, with atomic-skill segmentation and quality annotations.
- MolmoAct2-BimanualYAM: over 720 hours of real bimanual manipulation with multi-view observations and language annotations.
- HABIT: over 10K episodes, 160 hours, 60 tasks involving human-present manipulation.
- ActionNet: over 30K teleoperated trajectories, about 140 hours, dexterous bimanual humanoid manipulation.
Before filtering, robot data totals 13,867.71 hours across 1,370,003 episodes. After filtering, the authors retain 11,302.20 hours across 1,247,656 episodes, corresponding to 1,101.047 million frames.
Robot data processing rules
The robot processing pipeline applies these stages in order:
- Signal anomaly and consistency filtering: removes non-finite values, abrupt outliers, and state-action inconsistencies; uses local cross-correlation alignment with a default agreement threshold of 0.65.
- Static boundary trimming: adaptively removes prolonged inactivity at episode boundaries while preserving intermediate pauses.
- Visual quality filtering: detects black frames, blur, compression artifacts, extreme temporal changes, blockiness, and saturation anomalies.
- Action magnitude guard: rejects end-effector actions with single-step translation over 0.2 m or rotation over 0.5 rad.
- Instruction correctness filtering: uses an LLM to remove empty, malformed, incomplete, or unclear instructions.
- Video-instruction consistency filtering: uses a vision-language model to compare gripper-event clips and global frames with the instruction.
- Manual semantic verification: sampled episodes are inspected after decoding joint trajectories and visualizing end-effector poses against recorded videos.
Egocentric and ego-to-robot data
The authors use:
- EgoDex: 829 hours of egocentric human demonstrations across 194 tabletop tasks, collected with Apple Vision Pro. It provides 30 Hz RGB video, 3D head/upper-body/hand poses, camera intrinsics, camera poses, and language descriptions.
- EgoVerse: an expanded snapshot of about 473K recordings, yielding 1.48 million segmented episodes and 3,819 hours. It includes egocentric video, 3D hand keypoints, 6-DoF head poses, and task descriptions, with subtask-level language annotations for industry-contributed data.
After filtering, EgoDex and EgoVerse together retain 4,061.35 hours from 4,640.04 hours, and 1,622,756 episodes from 1,818,396 episodes. Retained durations are 772.16 hours for EgoDex and 3,289.19 hours for EgoVerse, totaling 438.626 million frames.
Egocentric processing includes:
- Action alignment: converts human wrist poses into camera-relative end-effector states. Gripper openness is estimated from fingertip geometry, with 5 cm mapped to fully closed and 7 cm mapped to fully open.
- Action speed alignment: EgoDex and EgoVerse trajectories are slowed down by a factor of two.
- Egocentric data undergoes action alignment and action speed alignment only.
Ego2Robot conversion adds:
- Kinematic alignment: base selection is coupled with inverse kinematics, using coarse-to-fine trajectory validation to detect tracking failures and discontinuous IK branches.
- Visual alignment: SAM3 segments human regions, ProPainter removes them, and a rendered robot is composited into the inpainted scene.
- Depth-aware compositing uses metric hand keypoints and object masks, with validity masks retained for downstream training.
From filtered EgoDex and EgoVerse video, the pipeline covers 1,101.55 hours from 283,450 unique source episodes. It yields 5,633.77 robot-hours across 3,730,513 training episodes, totaling 608.262 million frames. Robot-hours are reported before the factor-of-two action speed slowdown.
UMI data
The authors use the public Hy-UMI-10K release:
- Approximately 250K episodes spanning 2,162 hours.
- Head-camera and wrist-camera RGB video paired with 6-DoF gripper poses, gripper openness, and task descriptions.
UMI filtering includes:
- Removing segments with quaternion norm ∣∣q∣∣2≤10−8.
- Removing segments where both hands are static: position variance below 5×10−4 m2 and rotation variance below 0.1 rad2.
- Applying trajectory anomaly filtering and rejecting excessive speed or abrupt position/orientation jumps.
After filtering, Hy-UMI-10K is reduced from 2,161.74 hours and 250,135 episodes to 2,075.42 hours and 224.146 million frames. Because retained UMI trajectories are segmented into multiple training episodes, the episode count rises to 416,053.
Usage in the model
The paper uses the filtered and converted data for joint training under the shared canonical representation. The robot data, ego data, ego-to-robot converted robot data, and UMI data all provide aligned state and action supervision. Exported ego-to-robot training segments must contain at least 32 action steps. The provided dataset section reports filtering and conversion volumes but does not specify the final training mixture ratios.
Method
The authors propose InternW0-∆, a directed world-action architecture that integrates pretrained visual dynamics, task-conditioned scene semantics, and action generation. As shown in the figure below, the framework couples a pretrained video expert and an action expert through a directed Mixture-of-Transformers (MoT). A frozen vision-language model provides scene-grounded task semantics to the action expert, while the video expert processes a lightweight sparse memory of anchor, recent, and current observations. Proprioceptive states condition both experts independently.
To handle heterogeneous camera configurations, multi-view observations are resized and arranged into a unified visual canvas, then encoded by a pretrained video VAE to produce visual latents. The model maintains a sparse visual memory comprising an episode-level global memory (anchor), a short-term memory from the previous action chunk (recent), and the current observation. This design preserves both long-term task context and short-term motion continuity without the computational overhead of a dense history. Proprioceptive states are projected into separate embedding spaces for the video and action experts via linear layers. For task conditioning, the video expert utilizes T5 embeddings of the language instruction to preserve pretrained language priors. To complement this with scene-level understanding, a frozen VLM jointly processes valid current views and the instruction, providing the action expert with rich multimodal tokens regarding the current environment and task.
To learn action-relevant dynamic representations without exposing realized future observations to the action prediction path, the authors introduce Causal Imprint tokens into the self-attention layers of the video expert. These tokens aggregate recent and current visual features to capture motion and state transitions. During training, they are supervised by adjacent differences between clean ground-truth video latents and aligned with intermediate features of the video expert. As shown in the figure below, the dual supervision mechanism for Causal Imprint combines direct regression on latent differences with feature alignment from future slices.
Furthermore, to enrich the video expert with geometric and motion-aware representations, a training-only distillation objective is employed using a frozen Track4World teacher. As depicted in the figure below, the teacher processes ground-truth video windows to generate 4D-aware descriptors containing geometry, 2D/3D motion, and camera statistics. A student branch extracts clean-condition hidden tokens from the video expert, aggregates them via a Transformer decoder, and aligns the resulting descriptor with the cached teacher descriptor using an auxiliary MSE loss. This distillation transfers geometric priors without adding inference cost.
The video and action experts retain separate modality-specific parameters but exchange information through masked joint attention within the MoT blocks. At each coupled transformer block, both experts independently compute query, key, and value projections, which are concatenated for a single attention operation before being routed back to their respective streams. As shown in the figure below, the token types, the MoT layer structure, and the attention mask are detailed. The mask strictly controls information flow: future-frame tokens can attend to observed context for future prediction, but Causal Imprint and action tokens are blocked from directly accessing realized future frames. This directed flow ensures the action expert conditions only on observed context and learned predictive representations.
The model is jointly optimized using flow matching objectives for both future video generation and action prediction. The Causal Imprint tokens are trained with a squared-error loss against adjacent latent differences and a cosine alignment loss against detached future-video features. The 4D-aware distillation uses an auxiliary MSE loss to align the student descriptor with the cached teacher descriptor. The overall objective combines these losses with weighted coefficients.
At inference time, InternW0- generates action chunks without sampling or decoding future video. As illustrated in the figure below, the process begins with a single video-expert prefill where anchor, recent, and current observations are encoded and processed alongside Causal Imprint tokens. The self-attention key/value projections of these visual tokens are cached at each MoT layer. Subsequently, the action expert performs iterative action-only denoising over N flow steps, reusing the cached visual features and attending only to the action stream. This design eliminates the need for future-video rollout, requiring only one video-expert prefill and N action-expert updates per observation cycle.
Experiment
InternW0-∆ is evaluated on four simulation benchmarks (LIBERO-Plus, RoboTwin 2.0, EBench, and RoboDojo) and eleven real-robot tasks across four platforms, with additional experiments covering training efficiency, model components, data sources, and GPT-guided correction. The training optimizations yield end-to-end throughput speedups of 2.11x to 3.02x, while the model achieves the best overall results on LIBERO-Plus and RoboTwin 2.0, strong long-horizon and perturbation performance, and real-robot gains from 20% and 0% to 95% after pretraining, including generalization to unseen water levels and cross-embodiment deployment. Ablations attribute the improvements to sparse memory context, the VLM representation, causal imprint with representation alignment, 4D-aware distillation, robot-aligned human data, and GPT-based execution correction, which substantially improves difficult RoboDojo tasks.
The canonical action layout assigns fixed semantic slots to robot actions across an 80-dimensional vector, covering dual-arm manipulation and body mobility. Paired left and right arm components receive separate index ranges for joints, end-effector, gripper, and hand, while torso joints occupy a dedicated block. This format standardizes heterogeneous action dimensions for cross-embodiment learning. Dual-arm slots are separated by side, with each side sharing the same structure for arm joints, end-effector, gripper, and hand components. Torso joints are assigned a dedicated body mobility range outside the arm-specific slots.
Filtering reduces the robot collection by about 18 percent of hours and 9 percent of episodes, while retaining roughly 1.10 billion frames. Egocentric human data keeps approximately 88 percent of hours and 89 percent of episodes, and UMI data retains nearly all of its duration after filtering. Robot data has the largest absolute hour reduction during filtering, losing about 2,566 hours and roughly 122,000 episodes. UMI filtering removes only a small share of duration, and trajectory segmentation raises the episode count from about 250,000 to over 416,000.
The Ego2Robot pipeline converts filtered EgoDex and EgoVerse egocentric video into robot training data. Overall, roughly a quarter of source video hours and less than a fifth of source episodes complete synthesis successfully. The exported dataset contains substantially more robot-hours and episodes than the converted source coverage, with EgoVerse contributing most of the final training volume. EgoVerse provides the majority of source hours, converted hours, final robot-hours, and frames, while EgoDex converts a larger share of its input hours and episodes. Successful conversion covers about one quarter of source video hours and under one fifth of source episodes, but the final exported robot episodes are far more numerous than covered source episodes. The final robot dataset expands to millions of episodes and hundreds of millions of frames, with more than five robot-hours produced per successfully converted source hour.
Pretraining uses a heterogeneous hybrid dataset led by real robot demonstrations, with smaller shares of ego-to-robot, UMI-collected, and egocentric demonstrations. All sources are standardized to a canonical 80-dimensional state and action space, and training samples pair 33 video frames with 32 action steps at a video-to-action frequency ratio of 4. Under this configuration, raw egocentric video gives limited downstream success-rate gains, although this may not reflect representation quality or larger-scale benefits. The hybrid pretraining mixture is dominated by real robot demonstrations, with smaller contributions from ego-to-robot, UMI, and egocentric sources. Action loss weights are highest for robot demonstrations, lower for UMI data, and lower still for ego-to-robot and egocentric demonstrations.
The four simulation post-training configurations share a common 384x256 packed multi-view canvas and a consistent trajectory window and action chunking setup. Most benchmarks use three views, z-score normalization, and 10 epochs. LIBERO-Plus is the exception with two views, min-max normalization, and 15 epochs, while RoboTwin 2.0 is trained on clean data and evaluated on randomized data. All four benchmarks share a 384x256 packed multi-view resolution. EBench, RoboDojo, and RoboTwin 2.0 align on three views, z-score normalization, and 10 epochs. LIBERO-Plus differs by using two views, min-max normalization, and 15 epochs, while RoboTwin 2.0 is trained on the clean split and evaluated on the randomized split.
The evaluation protocol standardizes heterogeneous robot, egocentric, and UMI data into a canonical 80-dimensional action space, filters the collections, and converts egocentric video into robot training trajectories through the Ego2Robot pipeline. Filtering retains most egocentric and UMI data while removing a notable share of robot hours, and the conversion process shows that successfully covering only a minority of source egocentric video still produces a large final robot dataset. Pretraining uses a hybrid mixture dominated by real robot demonstrations, with lower action loss weights for UMI and ego-to-robot data, and raw egocentric video provides limited downstream success-rate gains. The simulation post-training benchmarks share a packed multi-view input and action chunking setup, with LIBERO-Plus and RoboTwin 2.0 differing mainly in view count, normalization, epochs, and clean-to-randomized evaluation protocol.