Command Palette
Search for a command to run...
GPU加速による宇宙機ランデブ・近傍運用のための天体力学ワールドモデル
GPU加速による宇宙機ランデブ・近傍運用のための天体力学ワールドモデル
Duncan Eddy Isaac R. Ward Grace Ra Kim Mykel J. Kochenderfer
概要
ワールドモデルは表現学習における新興パラダイムであり、エージェントがオフラインの軌跡データから状態-行動ダイナミクスと観測モデルを共同で学習し、不確実性推定を伴う多段階計画と軌跡予測を可能にする。ロボティクスやゲーム環境では強力な結果が示されているが、我々の知る限り、宇宙領域にはこれまで適用されていない。本論文では、協力的および非協力的な宇宙機のランデブ・近傍運用に対するワールドモデルベースのアプローチを紹介し、3つの貢献を行う。第一に、国際宇宙ステーション(ISS)へのドッキングシミュレーション環境をオープンソースとして導入する。これはJAXベースで、宇宙機の軌道および姿勢ダイナミクスの並列GPUシミュレーションをサポートし、ワールドモデルのトレーニングに必要な数千の状態-行動遷移を生成する。第二に、Out-of-this-World-Modelを導入する。これはトランスフォーマーベースのワールドモデルで、相対運動学的状態と機体固定カメラ画像を潜在状態にエンコードし、指令された推力と制御トルクの下でのその進化を一段フローマッチングを用いて予測する。このモデルは将来の観測に対する分布を生成し、確率論的ダイナミクスを捉え、各タイムステップの不確実性推定を提供し、DreamerV3スタイルの事後補正ベースラインよりも少ない訓練可能パラメータとハイパーパラメータで高い予測性能を達成する。第三に、このアプローチをキープアウトゾーン制約下でISSに自動ドッキングするカプセルに適用し、強化学習ベースラインと比較してサンプル効率とタスク性能が向上すること(ポート全体でのドッキング成功率53%対29%)、分布外汎化が大幅に優れること(未見のポートではワールドモデルの成功率がベースラインの2倍以上、40%対17%)、およびドッキング接近中に遭遇する異常物体を98%の分類精度で検出することを実証する。我々は、このパラダイムのさらなる研究を可能にするために、シミュレーション環境とモデルアーキテクチャをオープンソース化する。
One-sentence Summary
Researchers from Stanford University introduce Out-of-this-World-Model, a JAX-based transformer world model for spacecraft rendezvous and proximity operations that encodes relative kinematics and camera imagery into a latent state, predicts its evolution under commanded thrusts via one-step flow matching, and achieves 53% versus 29% docking success, more than doubles out-of-distribution generalization on held-out ports (40% versus 17%), and attains 98% anomaly detection accuracy in an ISS docking environment.
Key Contributions
- Introduces AstroJAX, an open-source, JAX-based simulation environment for International Space Station docking that runs spacecraft orbit and attitude dynamics in parallel on GPUs, enabling generation of 500,000 state-action transitions for world model training.
- Presents Out-of-this-World-Model, a transformer-based world model that fuses relative kinematic states and body-fixed camera imagery into a latent representation and predicts future observations via one-step flow matching, achieving greater predictive performance than DreamerV3-style baselines with fewer trainable parameters and hyperparameters.
- Applying the approach to a capsule docking with the ISS under keep-out-zone constraints demonstrates improved docking success over reinforcement learning baselines (53% versus 29% across ports), more than doubles success on held-out ports (40% versus 17%), and detects anomalous objects during approach with 98% classification accuracy.
Introduction
Spacecraft rendezvous and proximity operations (RPO) are shifting from rare, human-supervised missions to routine tasks across civil, military, and commercial sectors, with applications like docking at the ISS, on-orbit servicing, and life extension of satellites. Traditional guidance, navigation, and control (GNC) systems, which rely on Kalman filters and model predictive control, struggle with rich sensor data like camera imagery, requiring preprocessing pipelines and assuming accurate dynamics models. Existing learning-based methods, such as direct regression or physics-informed neural networks, fail to predict sensor observations or degrade over long horizons, while reinforcement learning offers reactive policies without reusable system models. The authors introduce a world model approach that learns joint state-action dynamics and observation models from trajectory data, presenting AstroJAX, a GPU-accelerated simulation framework, and Out-of-this-World-Model (OWM), a transformer-based architecture that fuses kinematic and visual data to predict future observations. This model demonstrates improved docking success across ISS ports, generalizes to unseen berthing ports, and detects anomalies with 98% accuracy, marking the first application of world models to space operations.
Dataset
Dataset Composition and Sources
The authors construct a dataset for learning world models in an ISS docking environment. The data is generated synthetically from three simulation environments that share a common task layer but differ in equation-of-motion fidelity. The paper's results use the highest-fidelity member, iss-numerical, a full numerical simulation of a Dragon-class chaser and the ISS chief.
The simulation propagates both vehicles as independent state vectors in the Earth-centered inertial frame, incorporating zonal gravity harmonics through degree four, third-body accelerations from the Sun and Moon, and atmospheric drag with Harris-Priester density. The chaser's attitude uses quaternion kinematics and rigid-body Euler dynamics. The full state is 21-dimensional and integrates with a fourth-order Runge-Kutta scheme at a 0.05 s timestep, with episodes lasting at most 7200 steps (360 s).
Key Details for Each Subset
The dataset comprises trajectories generated by three behavior policies, mixed in a 0.30/0.35/0.35 ratio for the training split:
- Random policy: Samples uniform force and torque commands, producing undirected drift that covers the state space far from the station.
- Orbit policy: Commands a proportional-derivative-tracked circumnavigation at a sampled radius between 80 m and 130 m, exposing the model to sustained lateral motion and varied viewing geometries.
- Dock policy: Flies a critically damped proportional-derivative approach to a sampled docking port. About half of these approaches meet contact conditions; the remainder end in recorded collisions, covering contact-adjacent failure modes.
Each episode initializes the chaser at a uniformly sampled radius between 100 m and 225 m from the station origin, with the nose pointed at the station and an epoch drawn uniformly from a seven-day window, which sweeps the solar beta angle and lighting conditions.
Training Split and Held-Out Ports
The training split's dock lane targets five of eight ports: Harmony forward, Harmony nadir, Zvezda aft, Pirs nadir, and Rassvet nadir. The three remaining ports (Harmony zenith, Poisk zenith, Unity nadir) appear only in evaluation, providing a held-out generalization test with unseen goal poses, approach corridors, and visual context.
Sensor Noise and Observation Models
Each frame records the true state, a noisy observation vector, the action, the reward, and a rendered first-person camera view, packaged in LeRobot format with normalization statistics computed on the training split alone.
Sensor noise follows one of three models:
- Cooperative: Differential-GNSS-class relative navigation with a fixed position error budget.
- Non-cooperative: Vision-based navigation where position error grows with range and velocity estimates are coarser.
- No noise: Enables isolation and quantification of world-model reconstruction errors.
Attitude and rate noise, assumed to come from the chaser's own star tracker and gyroscopes, are identical across the two noise presets. All noise magnitudes are total root-mean-square errors.
Processing and Rendering Details
The kinematic observation channel reports a 13-dimensional relative view (relative position, velocity, attitude quaternion, and angular velocity) through a configurable Gaussian sensor model. The visual channel is a 512 x 512 first-person view from a camera on the chaser's nose, with an 82-degree vertical field of view. Frames are rendered from recorded simulation states at the 20 Hz simulation rate using a physically based renderer that includes a full-globe textured Earth, starfield, Moon, and epoch-dependent Sun lighting.
Data Usage in the Model
The behavior policies act on noisy measurements rather than true states, so recorded action-outcome pairs reflect the aleatoric transition uncertainty of acting on an observed state. This is intentional, as the world model never trains on the reward function. The training data is used to train the world model to predict future states and observations, with the mixed-policy corpus providing broad coverage of goal-directed, undirected, and contact-adjacent motion.
Method
The authors formulate the docking scenario as a discrete-time partially observable decision process, where the world model approximates the distribution over the next observation conditioned on a history of past observations and actions. The proposed architecture features a modality-parameterized design, allowing the same backbone to process any combination of input streams.
As shown in the figure below:
Each input stream contributes a fixed number of tokens to a per-timestep token vector. A lightweight multilayer perceptron projects the kinematic state measurements, while a vision transformer encodes each camera frame into image tokens through a learned attention bottleneck that cross-attends a small set of latent queries against the patch embeddings of the frame. A second multilayer perceptron projects the action into an additional token, which is appended to the per-timestep vector as conditioning. The concatenation of these state, image, and action tokens forms the latent state for each timestep.
The backbone is a factorized space-time transformer. Each block applies spatial attention, which performs bidirectional attention among the tokens of a single timestep to fuse state, action, and image information. This is followed by temporal attention, which attends causally across timesteps within a sliding window using rotary position embeddings. Factoring the attention in this manner reduces the computational cost from quadratic in the full token sequence to quadratic in each axis separately. Furthermore, a key-value cache enables autoregressive rollouts that are linear in horizon length, addressing throughput bottlenecks during planning.
Prediction is performed entirely in latent space using a flow-matching head applied independently to each token. Conditioned on the backbone output for the current timestep, the head learns a velocity field that transports a standard Gaussian sample to the residual between the next latent state and the current one. The sampled latent is appended to the token history, and the model rolls forward via latent-space autoregression without decoding to observations during the loop. Modality-specific decoder heads map latents back to predicted images and states only when observation-space output is required.
The training procedure operates on windows of context steps followed by rollout steps sampled from recorded trajectories. The flow head is trained with a rectified flow objective on the latent residual. For a noise sample ε∼N(0,I) and a noise level τ∼U(0,1), the noised input lies on the straight path between data and noise:
xτ=(1−τ)Δzt+1+τεand the head minimizes the squared error to regress the constant velocity of that path:
Lflow=∥fθ(xτ,τ,ht)−(ε−Δzt+1)∥2To ground the latent space in observations, the authors incorporate two families of reconstruction terms. A per-modality decode loss reconstructs each observation stream from the predicted latent:
Lmdec=∣S∣1t∈S∑∥gm(z^t)−otm∥2while a latent round-trip anchor reconstructs the true observation from its own encoding to keep the token codec close to an identity map:
Lmrt=T1t∑∥gm(zt)−otm∥2These terms combine as a weighted sum:
L=λflowLflow+m∑wmLmdec+m∑βmLmrtDuring training, the model rolls forward autoregressively on its own predictions, with a teacher-forcing probability annealed from one to zero over the initial epochs. The optimization utilizes AdamW with mixed bfloat16 precision and gradient-norm clipping.
Experiment
The experiments evaluate a world model based planner for ISS docking, comparing it to a PPO reinforcement learning baseline across sensor noise regimes on both training and held-out ports. The flow-matching world model achieves higher prediction quality than a Dreamer style baseline, and its planner matches or outperforms PPO on in-distribution docks while generalizing to all three unseen ports, where PPO fails on two. However, PPO achieves stricter terminal contact conditions due to more reward tuning. The world model also detects a novel anomalous object with 98% accuracy, concentrating predictive uncertainty on the unexpected element.
AstroJAX is an open-source, differentiable astrodynamics library implemented in JAX, covering gravity models, perturbations, ephemerides, attitude dynamics, relative motion, frames, time, and propagation. Validation against the brahe library shows close agreement for most models, with the exception of the NRLMSISE-00 atmospheric density model, which currently has limited accuracy. The library includes a wide range of models, from point-mass and spherical harmonic gravity to atmospheric drag, SRP, and multiple integrators. Frame transformations match IAU SOFA reference values to high precision, and most accelerations agree with brahe within relative tolerances. The NRLMSISE-00 atmospheric density model is the main outlier, with agreement bounded near 15%.
The generated datasets include a training split of 500,000 transitions with a mixed policy ratio and a validation split of 50,000 transitions using only the dock policy. Training covers five docking ports, while validation covers all eight, with three ports held out for generalization testing. Each frame stores state, noisy observations, action, reward, and rendered camera view, with noise models applied to actions to reflect aleatoric uncertainty. Training mixes random, orbit, and dock policies in a 0.30/0.35/0.35 ratio, while validation uses only the dock policy. The training split uses five docking ports, leaving three ports (Harmony zenith, Poisk zenith, Unity nadir) unseen during training for held-out evaluation. Sensor noise is applied to the observations used by the behavior policies, making the recorded transition dynamics reflect real-world uncertainty rather than deterministic true-state dynamics.
Sensor noise is modeled under three presets: no noise, cooperative, and non-cooperative. Cooperative noise uses a fixed position error, while non-cooperative position error scales with range. Attitude and body-rate noise are identical across cooperative and non-cooperative presets. Non-cooperative position error grows with range at 1% of range, whereas cooperative uses a fixed 0.05 m error. Velocity noise is larger for non-cooperative (0.03 m/s) than cooperative (0.002 m/s). Attitude and body-rate noise are the same for both cooperative and non-cooperative models.
The table specifies the reward shaping parameters used to train a proximal policy optimization (PPO) baseline for docking. The weights and shaping kernels are tuned to prioritize slow, pointed approaches, with a large penalty on position error and a moderate penalty on attitude error, while providing small positive rewards for proximity and progress. These parameters were necessary because the default environment configuration made temporal-difference learning ineffective. The position error penalty dominates the reward with a weight of -0.9 and a wide shaping scale, encouraging the agent to reduce distance to the dock. Attitude and body rate penalties are relatively small but activate only when the agent is close to the dock, shaping pointed final approaches. A progress reward of -2.0 on the change in distance relative to the position scale encourages steady advancement toward the goal. Positive proximity and alignment bonuses are active only within a small radius, providing sparse additional incentive near the dock. The baseline required substantial tuning of these parameters to achieve any meaningful docking behavior, highlighting the difficulty of the task for model-free RL.
The table documents a PPO baseline configuration adapted for a docking task, including 25M total environment steps, 32 parallel environments, and a 0.997 discount factor. The baseline uses a 256x3 network with a 1 Hz control step to address reward sparsity, and its reward function incorporates multiple shaping terms. The cited analysis shows the planner outperforms PPO under position-only success, but PPO achieves higher success under full contact conditions, while the planner exhibits zero escape rate compared to PPO's high escape rate on zenith ports. PPO is trained with 25M steps across 32 parallel environments and a 0.997 discount factor. The baseline uses a 1 Hz control step to mitigate reward sparsity, as the native 20 Hz step renders the terminal bonus negligible. Under position-only success, the planner outperforms PPO at all tolerances, but under full contact conditions PPO achieves higher success rates, attributed to tuning effort. The planner has a 0% escape rate across all ports, while PPO escapes in 88% of rollouts at each zenith port, indicating PPO's lower collision rate reflects missed approaches. The planner's residual failure mode is collisions on zenith approaches, whereas PPO's is losing acquisition via escapes.
AstroJAX, a differentiable astrodynamics library in JAX, was validated against brahe, showing close agreement for most models, except the NRLMSISE-00 atmospheric density model, whose accuracy is bounded near 15%; frame transformations match IAU SOFA references to high precision. A docking RL dataset was generated with separate training and validation splits, using 5 training ports and 3 held-out ports, and sensor noise was modeled under three presets (none, cooperative, non-cooperative), with non-cooperative errors scaling with range. For a PPO baseline, reward shaping was heavily tuned to address reward sparsity, with a dominant position-error penalty and near-dock bonuses; the planner outperformed PPO under position-only success, but PPO achieved higher success under full contact, while the planner had a 0% escape rate versus PPO's 88% escape rate on zenith ports, indicating PPO's collisions were more often missed approaches.