HyperAIHyperAI

Command Palette

Search for a command to run...

نماذج العالم الفضائي المعجَّلة بوحدات معالجة الرسوميات لعمليات الالتقاء والقرب للمركبات الفضائية

Duncan Eddy Isaac R. Ward Grace Ra Kim Mykel J. Kochenderfer

الملخص

تُعد نماذج العالم范式ًا ناشئًا في تعلم التمثيل، حيث يتعلم الوكيل بشكل مشترك ديناميكيات الحالة-الفعل ونماذج الملاحظة من بيانات المسارات غير المتصلة بالإنترنت، مما يتيح التخطيط متعدد الخطوات والتنبؤ بالمسارات مع تقديرات لعدم اليقين. وقد أظهرت هذه النماذج نتائج قوية في بيئات الروبوتات والألعاب، ولكن، على حد علمنا، لم تُطبَّق سابقًا على مجال الفضاء. تقدم هذه الورقة نهجًا قائمًا على نماذج العالم لعمليات الالتقاء والقرب للمركبات الفضائية التعاونية وغير التعاونية، وتُسهم بثلاثة إسهامات. أولاً، نقدم بيئة محاكاة مفتوحة المصدر قائمة على JAX لالتحام محطة الفضاء الدولية (ISS) تدعم المحاكاة المتوازية لوحدات معالجة الرسوميات لديناميكيات مدار المركبة الفضائية واتجاهها، مما يولد آلاف انتقالات الحالة-الفعل التي يتطلبها تدريب نماذج العالم. ثانيًا، نقدم "Out-of-this-World-Model"، وهو نموذج عالم قائم على المحولات يرمّز الحالات الحركية النسبية وصور الكاميرا المثبتة على الجسم في حالة كامنة ويتنبأ بتطورها تحت الدفعات المأمورة وعزوم التحكم باستخدام مطابقة التدفق أحادية الخطوة. ينتج النموذج توزيعًا على الملاحظات المستقبلية، ملتقطًا الديناميكيات العشوائية ويوفر تقديرات لعدم اليقين لكل خطوة زمنية، ويحقق أداءً تنبؤيًا أكبر من خطوط الأساس لتصحيح الخلفية بأسلوب DreamerV3 مع معلمات قابلة للتدريب وفرط معلمات أقل. ثالثًا، نطبق النهج على كبسولة تلتحم تلقائيًا بمحطة الفضاء الدولية تحت قيود منطقة الحظر، مما يُظهر كفاءة عينات وأداء مهام محسّنًا مقارنة بخطوط الأساس للتعلم المعزز (53% مقابل 29% من نجاح الالتحام عبر المنافذ)، وتعميمًا أفضل بشكل كبير خارج التوزيع (على المنافذ المحجوزة، يضاعف نموذج العالم نجاح خط الأساس بأكثر من الضعف، 40% مقابل 17%)، واكتشاف الأجسام الشاذة التي تُواجه أثناء نهج الالتحام بدقة تصنيف 98%. نتيح بيئة المحاكاة وهندسة النموذج كمصدر مفتوح لتمكين مزيد من الدراسة لهذا النهج.

One-sentence Summary

Researchers from Stanford University introduce Out-of-this-World-Model, a JAX-based transformer world model for spacecraft rendezvous and proximity operations that encodes relative kinematics and camera imagery into a latent state, predicts its evolution under commanded thrusts via one-step flow matching, and achieves 53% versus 29%53\% \text{ versus } 29\%53% versus 29% docking success, more than doubles out-of-distribution generalization on held-out ports (40% versus 17%40\% \text{ versus } 17\%40% versus 17%), and attains 98%98\%98% anomaly detection accuracy in an ISS docking environment.

Key Contributions

  • Introduces AstroJAX, an open-source, JAX-based simulation environment for International Space Station docking that runs spacecraft orbit and attitude dynamics in parallel on GPUs, enabling generation of 500,000 state-action transitions for world model training.
  • Presents Out-of-this-World-Model, a transformer-based world model that fuses relative kinematic states and body-fixed camera imagery into a latent representation and predicts future observations via one-step flow matching, achieving greater predictive performance than DreamerV3-style baselines with fewer trainable parameters and hyperparameters.
  • Applying the approach to a capsule docking with the ISS under keep-out-zone constraints demonstrates improved docking success over reinforcement learning baselines (53% versus 29% across ports), more than doubles success on held-out ports (40% versus 17%), and detects anomalous objects during approach with 98% classification accuracy.

Introduction

Spacecraft rendezvous and proximity operations (RPO) are shifting from rare, human-supervised missions to routine tasks across civil, military, and commercial sectors, with applications like docking at the ISS, on-orbit servicing, and life extension of satellites. Traditional guidance, navigation, and control (GNC) systems, which rely on Kalman filters and model predictive control, struggle with rich sensor data like camera imagery, requiring preprocessing pipelines and assuming accurate dynamics models. Existing learning-based methods, such as direct regression or physics-informed neural networks, fail to predict sensor observations or degrade over long horizons, while reinforcement learning offers reactive policies without reusable system models. The authors introduce a world model approach that learns joint state-action dynamics and observation models from trajectory data, presenting AstroJAX, a GPU-accelerated simulation framework, and Out-of-this-World-Model (OWM), a transformer-based architecture that fuses kinematic and visual data to predict future observations. This model demonstrates improved docking success across ISS ports, generalizes to unseen berthing ports, and detects anomalies with 98% accuracy, marking the first application of world models to space operations.

Dataset

Dataset Composition and Sources

The authors construct a dataset for learning world models in an ISS docking environment. The data is generated synthetically from three simulation environments that share a common task layer but differ in equation-of-motion fidelity. The paper's results use the highest-fidelity member, iss-numerical, a full numerical simulation of a Dragon-class chaser and the ISS chief.

The simulation propagates both vehicles as independent state vectors in the Earth-centered inertial frame, incorporating zonal gravity harmonics through degree four, third-body accelerations from the Sun and Moon, and atmospheric drag with Harris-Priester density. The chaser's attitude uses quaternion kinematics and rigid-body Euler dynamics. The full state is 21-dimensional and integrates with a fourth-order Runge-Kutta scheme at a 0.05 s timestep, with episodes lasting at most 7200 steps (360 s).

Key Details for Each Subset

The dataset comprises trajectories generated by three behavior policies, mixed in a 0.30/0.35/0.35 ratio for the training split:

  • Random policy: Samples uniform force and torque commands, producing undirected drift that covers the state space far from the station.
  • Orbit policy: Commands a proportional-derivative-tracked circumnavigation at a sampled radius between 80 m and 130 m, exposing the model to sustained lateral motion and varied viewing geometries.
  • Dock policy: Flies a critically damped proportional-derivative approach to a sampled docking port. About half of these approaches meet contact conditions; the remainder end in recorded collisions, covering contact-adjacent failure modes.

Each episode initializes the chaser at a uniformly sampled radius between 100 m and 225 m from the station origin, with the nose pointed at the station and an epoch drawn uniformly from a seven-day window, which sweeps the solar beta angle and lighting conditions.

Training Split and Held-Out Ports

The training split's dock lane targets five of eight ports: Harmony forward, Harmony nadir, Zvezda aft, Pirs nadir, and Rassvet nadir. The three remaining ports (Harmony zenith, Poisk zenith, Unity nadir) appear only in evaluation, providing a held-out generalization test with unseen goal poses, approach corridors, and visual context.

Sensor Noise and Observation Models

Each frame records the true state, a noisy observation vector, the action, the reward, and a rendered first-person camera view, packaged in LeRobot format with normalization statistics computed on the training split alone.

Sensor noise follows one of three models:

  • Cooperative: Differential-GNSS-class relative navigation with a fixed position error budget.
  • Non-cooperative: Vision-based navigation where position error grows with range and velocity estimates are coarser.
  • No noise: Enables isolation and quantification of world-model reconstruction errors.

Attitude and rate noise, assumed to come from the chaser's own star tracker and gyroscopes, are identical across the two noise presets. All noise magnitudes are total root-mean-square errors.

Processing and Rendering Details

The kinematic observation channel reports a 13-dimensional relative view (relative position, velocity, attitude quaternion, and angular velocity) through a configurable Gaussian sensor model. The visual channel is a 512 x 512 first-person view from a camera on the chaser's nose, with an 82-degree vertical field of view. Frames are rendered from recorded simulation states at the 20 Hz simulation rate using a physically based renderer that includes a full-globe textured Earth, starfield, Moon, and epoch-dependent Sun lighting.

Data Usage in the Model

The behavior policies act on noisy measurements rather than true states, so recorded action-outcome pairs reflect the aleatoric transition uncertainty of acting on an observed state. This is intentional, as the world model never trains on the reward function. The training data is used to train the world model to predict future states and observations, with the mixed-policy corpus providing broad coverage of goal-directed, undirected, and contact-adjacent motion.

Method

The authors formulate the docking scenario as a discrete-time partially observable decision process, where the world model approximates the distribution over the next observation conditioned on a history of past observations and actions. The proposed architecture features a modality-parameterized design, allowing the same backbone to process any combination of input streams.

As shown in the figure below:

Each input stream contributes a fixed number of tokens to a per-timestep token vector. A lightweight multilayer perceptron projects the kinematic state measurements, while a vision transformer encodes each camera frame into image tokens through a learned attention bottleneck that cross-attends a small set of latent queries against the patch embeddings of the frame. A second multilayer perceptron projects the action into an additional token, which is appended to the per-timestep vector as conditioning. The concatenation of these state, image, and action tokens forms the latent state for each timestep.

The backbone is a factorized space-time transformer. Each block applies spatial attention, which performs bidirectional attention among the tokens of a single timestep to fuse state, action, and image information. This is followed by temporal attention, which attends causally across timesteps within a sliding window using rotary position embeddings. Factoring the attention in this manner reduces the computational cost from quadratic in the full token sequence to quadratic in each axis separately. Furthermore, a key-value cache enables autoregressive rollouts that are linear in horizon length, addressing throughput bottlenecks during planning.

Prediction is performed entirely in latent space using a flow-matching head applied independently to each token. Conditioned on the backbone output for the current timestep, the head learns a velocity field that transports a standard Gaussian sample to the residual between the next latent state and the current one. The sampled latent is appended to the token history, and the model rolls forward via latent-space autoregression without decoding to observations during the loop. Modality-specific decoder heads map latents back to predicted images and states only when observation-space output is required.

The training procedure operates on windows of context steps followed by rollout steps sampled from recorded trajectories. The flow head is trained with a rectified flow objective on the latent residual. For a noise sample ε∼N(0,I)\varepsilon \sim \mathcal{N}(0, I)ε∼N(0,I) and a noise level τ∼U(0,1)\tau \sim \mathcal{U}(0, 1)τ∼U(0,1), the noised input lies on the straight path between data and noise:

xτ=(1−τ)Δzt+1+τε\mathbf{x}_{\tau} = (1 - \tau) \Delta \mathbf{z}_{t+1} + \tau \varepsilonxτ​=(1−τ)Δzt+1​+τε

and the head minimizes the squared error to regress the constant velocity of that path:

Lflow=∥fθ(xτ,τ,ht)−(ε−Δzt+1)∥2\mathcal{L}_{\mathrm{flow}} = \left\| f_{\boldsymbol{\theta}}(\mathbf{x}_{\tau}, \boldsymbol{\tau}, \mathbf{h}_{t}) - (\boldsymbol{\varepsilon} - \Delta \mathbf{z}_{t+1}) \right\|^{2}Lflow​=∥fθ​(xτ​,τ,ht​)−(ε−Δzt+1​)∥2

To ground the latent space in observations, the authors incorporate two families of reconstruction terms. A per-modality decode loss reconstructs each observation stream from the predicted latent:

Lmdec=1∣S∣∑t∈S∥gm(z^t)−otm∥2\mathcal{L}_{m}^{\mathrm{dec}} = \frac{1}{|\mathscr{S}|} \sum_{t \in \mathscr{S}} \left\| g_{m}(\hat{\mathbf{z}}_{t}) - \mathbf{o}_{t}^{m} \right\|^{2}Lmdec​=∣S∣1​t∈S∑​∥gm​(z^t​)−otm​∥2

while a latent round-trip anchor reconstructs the true observation from its own encoding to keep the token codec close to an identity map:

Lmrt=1T∑t∥gm(zt)−otm∥2\mathcal{L}_{m}^{\mathrm{rt}} = \frac{1}{T} \sum_{t} \left\| g_{m}(\mathbf{z}_{t}) - \mathbf{o}_{t}^{m} \right\|^{2}Lmrt​=T1​t∑​∥gm​(zt​)−otm​∥2

These terms combine as a weighted sum:

L=λflowLflow+∑mwmLmdec+∑mβmLmrt\mathcal{L} = \lambda_{\text{flow}} \mathcal{L}_{\text{flow}} + \sum_{m} w_{m} \mathcal{L}_{m}^{\mathrm{dec}} + \sum_{m} \beta_{m} \mathcal{L}_{m}^{\mathrm{rt}}L=λflow​Lflow​+m∑​wm​Lmdec​+m∑​βm​Lmrt​

During training, the model rolls forward autoregressively on its own predictions, with a teacher-forcing probability annealed from one to zero over the initial epochs. The optimization utilizes AdamW with mixed bfloat16 precision and gradient-norm clipping.

Experiment

The experiments evaluate a world model based planner for ISS docking, comparing it to a PPO reinforcement learning baseline across sensor noise regimes on both training and held-out ports. The flow-matching world model achieves higher prediction quality than a Dreamer style baseline, and its planner matches or outperforms PPO on in-distribution docks while generalizing to all three unseen ports, where PPO fails on two. However, PPO achieves stricter terminal contact conditions due to more reward tuning. The world model also detects a novel anomalous object with 98% accuracy, concentrating predictive uncertainty on the unexpected element.

AstroJAX is an open-source, differentiable astrodynamics library implemented in JAX, covering gravity models, perturbations, ephemerides, attitude dynamics, relative motion, frames, time, and propagation. Validation against the brahe library shows close agreement for most models, with the exception of the NRLMSISE-00 atmospheric density model, which currently has limited accuracy. The library includes a wide range of models, from point-mass and spherical harmonic gravity to atmospheric drag, SRP, and multiple integrators. Frame transformations match IAU SOFA reference values to high precision, and most accelerations agree with brahe within relative tolerances. The NRLMSISE-00 atmospheric density model is the main outlier, with agreement bounded near 15%.

The generated datasets include a training split of 500,000 transitions with a mixed policy ratio and a validation split of 50,000 transitions using only the dock policy. Training covers five docking ports, while validation covers all eight, with three ports held out for generalization testing. Each frame stores state, noisy observations, action, reward, and rendered camera view, with noise models applied to actions to reflect aleatoric uncertainty. Training mixes random, orbit, and dock policies in a 0.30/0.35/0.35 ratio, while validation uses only the dock policy. The training split uses five docking ports, leaving three ports (Harmony zenith, Poisk zenith, Unity nadir) unseen during training for held-out evaluation. Sensor noise is applied to the observations used by the behavior policies, making the recorded transition dynamics reflect real-world uncertainty rather than deterministic true-state dynamics.

Sensor noise is modeled under three presets: no noise, cooperative, and non-cooperative. Cooperative noise uses a fixed position error, while non-cooperative position error scales with range. Attitude and body-rate noise are identical across cooperative and non-cooperative presets. Non-cooperative position error grows with range at 1% of range, whereas cooperative uses a fixed 0.05 m error. Velocity noise is larger for non-cooperative (0.03 m/s) than cooperative (0.002 m/s). Attitude and body-rate noise are the same for both cooperative and non-cooperative models.

The table specifies the reward shaping parameters used to train a proximal policy optimization (PPO) baseline for docking. The weights and shaping kernels are tuned to prioritize slow, pointed approaches, with a large penalty on position error and a moderate penalty on attitude error, while providing small positive rewards for proximity and progress. These parameters were necessary because the default environment configuration made temporal-difference learning ineffective. The position error penalty dominates the reward with a weight of -0.9 and a wide shaping scale, encouraging the agent to reduce distance to the dock. Attitude and body rate penalties are relatively small but activate only when the agent is close to the dock, shaping pointed final approaches. A progress reward of -2.0 on the change in distance relative to the position scale encourages steady advancement toward the goal. Positive proximity and alignment bonuses are active only within a small radius, providing sparse additional incentive near the dock. The baseline required substantial tuning of these parameters to achieve any meaningful docking behavior, highlighting the difficulty of the task for model-free RL.

The table documents a PPO baseline configuration adapted for a docking task, including 25M total environment steps, 32 parallel environments, and a 0.997 discount factor. The baseline uses a 256x3 network with a 1 Hz control step to address reward sparsity, and its reward function incorporates multiple shaping terms. The cited analysis shows the planner outperforms PPO under position-only success, but PPO achieves higher success under full contact conditions, while the planner exhibits zero escape rate compared to PPO's high escape rate on zenith ports. PPO is trained with 25M steps across 32 parallel environments and a 0.997 discount factor. The baseline uses a 1 Hz control step to mitigate reward sparsity, as the native 20 Hz step renders the terminal bonus negligible. Under position-only success, the planner outperforms PPO at all tolerances, but under full contact conditions PPO achieves higher success rates, attributed to tuning effort. The planner has a 0% escape rate across all ports, while PPO escapes in 88% of rollouts at each zenith port, indicating PPO's lower collision rate reflects missed approaches. The planner's residual failure mode is collisions on zenith approaches, whereas PPO's is losing acquisition via escapes.

AstroJAX, a differentiable astrodynamics library in JAX, was validated against brahe, showing close agreement for most models, except the NRLMSISE-00 atmospheric density model, whose accuracy is bounded near 15%; frame transformations match IAU SOFA references to high precision. A docking RL dataset was generated with separate training and validation splits, using 5 training ports and 3 held-out ports, and sensor noise was modeled under three presets (none, cooperative, non-cooperative), with non-cooperative errors scaling with range. For a PPO baseline, reward shaping was heavily tuned to address reward sparsity, with a dominant position-error penalty and near-dock bonuses; the planner outperformed PPO under position-only success, but PPO achieved higher success under full contact, while the planner had a 0% escape rate versus PPO's 88% escape rate on zenith ports, indicating PPO's collisions were more often missed approaches.


بناء الذكاء الاصطناعي بالذكاء الاصطناعي

من الفكرة إلى الإطلاق — سرّع تطوير الذكاء الاصطناعي الخاص بك مع المساعدة البرمجية المجانية بالذكاء الاصطناعي، وبيئة جاهزة للاستخدام، وأفضل أسعار لوحدات معالجة الرسومات.

البرمجة التعاونية باستخدام الذكاء الاصطناعي
وحدات GPU جاهزة للعمل
أفضل الأسعار

HyperAI Newsletters

اشترك في آخر تحديثاتنا
سنرسل لك أحدث التحديثات الأسبوعية إلى بريدك الإلكتروني في الساعة التاسعة من صباح كل يوم اثنين
مدعوم بواسطة MailChimp