Command Palette
Search for a command to run...
AgentGarten: 進化するエージェントのためのコード世界
AgentGarten: 進化するエージェントのためのコード世界
概要
インタラクティブな仮想世界は、エージェントが探索と対話を通じて学習することを可能にする。エージェントが学習できる内容は、練習に用いる環境によって制約される。その環境は、一貫した状態・規則・動力学を備えた忠実なものであり、かつ実世界の視覚分布に従う観測を提供する現実的なものでなければならない。多様な世界にわたってこの両方を達成することは依然としてボトルネックである。我々はAgentGartenを導入する。これは、シミュレータおよびゲームエンジンと共有ニューラルレンダラを結合し、リアルタイムの対話環境を構築するフレームワークである。そのシミュレーションバックエンドは永続的な世界状態を維持し、プログラムで定義された対話規則を実行する一方、レンダラは共通インターフェースを通じて出力される構造化条件から視覚観測を生成する。ニューラルレンダラを構築するために、我々は事前学習済み動画モデルを幾何条件に適応させ、提案するAdversarial Forcingによって蒸留し、リアルタイム対話に向けて推論を最適化する。Adversarial Forcingは、正確なリプレイを通じて履歴のプレフィルを微分可能にし、後続予測に対する損失がレンダラによる過去の観測の符号化を更新するようにするとともに、実データによる敵対的教師信号を加えて視覚品質を向上させる。AgentGartenでは、エージェントは視覚観測を通じて世界を知覚し、リアルタイムで対話し、各ラウンドの経験をプレイブックへ蒸留して、後続のエージェントがそれを継承し改良することで向上する。我々の実証研究は学習効率の大幅な向上を示しており、エージェントはわずか4ラウンドで学習するのに対し、従来の強化学習の対応手法では数百万ラウンドを要する。新しい世界はコードとして記述し、同じインターフェースを通じてレンダリングできるため、環境はエージェントとともに数と難度の両面で拡張でき、対話経験を通じて進化し続けるエージェントへの一歩となる。
One-sentence Summary
Researchers at MirroS, Tsinghua University, and Peking University introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments, adapts a pretrained video model using Adversarial Forcing and real-data adversarial supervision, and enables agents to learn from just four rounds compared with millions for a conventional reinforcement learning counterpart.
Key Contributions
- AgentGarten couples simulator or game-engine backends with a shared neural renderer to build real-time interactive environments; the backends maintain persistent world state and program-defined interaction rules, while the renderer converts structured conditions exported through a common interface into visual observations.
- Adversarial Forcing adapts a pretrained video model to geometry conditions, makes history prefilling differentiable through exact replay so later prediction losses update how prior observations are encoded, and adds real-data adversarial supervision to improve visual quality.
- The empirical study demonstrates that agents distill each round of experience into playbooks that later agents inherit and refine, learning from just 4 rounds compared with millions for a conventional reinforcement learning counterpart; in hide-and-seek, shelters appeared by round 4 and ramp crossings by round 10, and the same procedure improved outcomes in four further worlds.
Introduction
The authors target the need for scalable training environments in which agents learn from long trajectories of actions and visual observations. Such environments must preserve the consequences of past actions while allowing variation in layouts, objects, and task rules. Prior simulators and game engines offer explicit, inspectable state and programmable rules, but expanding their visual diversity is costly because it requires new 3D assets and rendering pipelines. Conversely, video world models synthesize rich observations, but their task-relevant state and transition rules remain implicit and hard to inspect, edit, or test. The authors introduce AgentGarten, a framework that combines programmable scene programs with a shared real-time neural renderer. Each scene runs in a simulator or game engine that maintains persistent state and executes interaction rules, while the renderer converts structured engine conditions into visual observations. To make this renderer interactive over extended rollouts, they adapt a pretrained bidirectional video model to block-causal generation and distill it with Adversarial Forcing, using exact replay for history gradients and an exact adversarial regularization scheme to preserve long-rollout visual quality.
Dataset
The authors describe a dataset of executable “code worlds” rather than a fixed set of images or videos. Each code world pairs an executable scene program with an appearance reference x0, and the scene program stores only coarse geometry needed for layout, silhouettes, occlusion, and motion dynamics.
-
Sources and composition
- Image-sourced code worlds: A single input image is used as both the appearance reference x0 and the layout reference. Perception and generation models extract instance masks, monocular depth, camera intrinsics, 3D bounding boxes, and object meshes.
- Text-sourced code worlds: The coding agent writes a scene program directly from geometric primitives and engine assets. An image editing model converts the initial render into x0 while preserving scene layout.
- Environment library: The collection includes navigation, manipulation, tool use, and multi-agent interaction scenes.
- Interactive browser games: The authors add bowling, a penalty shootout, and a crate-vault puzzle. Player input via keyboard, mouse, or gamepad advances the code world, while the rendered stream is the only view.
-
Processing and refinement
- For image-based scenes, the coding agent fits extracted assets to estimated geometry, completes unobserved regions as plausible spatial extensions, and binds physical properties and interaction rules.
- For text-based scenes, the initial render is transformed into x0 with layout-preserving image editing.
- Both modalities use an interactive refinement loop: the agent evaluates rollouts from multiple probe cameras, inspects engine verification queries for contact, collision clearance, and occlusion, and repairs faulty poses, unsupported structures, or ambiguous motions.
-
Schema and usage
- Each code world consists of an executable scene program and an appearance reference x0.
- Different scene programs and compatible engines form code worlds with a shared renderer.
- Because state evolution is explicit, recorded rollouts can be replayed and rendered again, including re-shooting from a different camera or in a different visual style without changing what happened.
-
Statistics and splits
- The provided excerpt does not report exact subset sizes, source dataset counts, filtering thresholds, training/evaluation split ratios, mixture ratios, or cropping details.
Method
The authors formulate a code world as a combination of a scene program, an execution engine, and a neural renderer. The program specifies the scene and interaction rules. Let st denote the scene state, πt the camera pose, and at the actions. The engine updates the state and captures structured conditions ct:
(st+1,πt+1)=fp(st,πt,at),ct=hp(st,πt)Given an appearance reference x0 and text description y, the neural renderer generates observations xt from visual history x<t:
xt=Rθ(x<t,c≤t,x0,y)This separation ensures that rendering errors do not accumulate in the state, allowing recorded state trajectories to be rendered under different appearances or cameras without altering the underlying events.
As shown in the interaction diagram below:
To construct these environments, a coding agent builds a code world from a single image or text description. When building from an image, the input serves simultaneously as the appearance reference and the reference layout. Perception and generation models lift the visible content into instance masks, monocular depth, camera intrinsics, and object meshes. The agent writes a scene program, fits these assets to the estimated geometry, and refines the scene through an interactive feedback loop. It evaluates test rollouts from multiple probe cameras and iteratively repairs faulty poses, unsupported structures, or ambiguous motions before deployment.
The process of building a code world from a single image is illustrated here:
The authors train a geometry-conditioned autoregressive video model to implement the neural renderer. The backbone separates understanding and generation towers. The frozen understanding tower encodes the caption into per-layer keys and values consumed by the generation tower. Geometry conditions, such as depth and surface normals, are encoded and concatenated with RGB tokens along the sequence dimension for joint attention. This preserves geometric alignment while allowing cross-token interactions without imposing a pixel-aligned inductive bias.
The final training stage, Adversarial Forcing, distills the teacher-forced model into a few-step renderer trained on its own rollouts. It combines distribution matching against the bidirectional teacher, exact history-gradient replay, and real-data adversarial training. To allow later losses to update the history-encoding computation without the memory cost of a fully differentiable rollout, the authors introduce exact replay. This method decouples rollout and gradient propagation using two passes. A no-gradient rollout records each block's input and clean output, and a differentiable replay recomputes both history and predictions with the specific visibility pattern.
The block-level attention masks for exact replay are depicted below:
For adversarial training, a trainable head aggregates intermediate features from the frozen teacher backbone into a discriminator logit. To regularize the discriminator with R1 and R2 penalties without requiring double backward through fused attention kernels, the authors compute the exact penalty and its exact head-parameter gradient. They utilize vector-Jacobian products to obtain gradients and Jacobian-vector products to carry directions forward through the frozen backbone, avoiding the materialization of the full Jacobian matrix.
The computation flow for exact R1/R2 regularization is shown below:
For continuous inference within the interaction loop, the renderer utilizes a bounded key-value cache. The cache maintains a permanent sink prefix, which includes the appearance reference, and a sliding window of recent history. Older frames are evicted as new blocks are denoised and published to the cache. This structure supports rollouts exceeding the training horizon through top-aligned rotary position remapping, preserving relative temporal distances while geometry tracks the true simulation timeline.
The structure of the inference cache is visualized below:
Experiment
The experiments evaluate visual quality, replay accuracy, and inference throughput on one NVIDIA H100 GPU with BF16 computation, 480×832 outputs, four-frame latent blocks, and a four-step sampler. Adversarial Forcing maintains natural textures and fine detail in long rollouts while Self Forcing produces repetitive patterns and loses detail, and block-by-block SDPA replay is bitwise identical and slightly faster than cached rollouts while FlexAttention replay introduces small error. The renderer delivers real-time 480×832 video generation, and agent experiments demonstrate that closed-loop grounding from neural-rendered first-person observations enables hide-and-seek agents to develop shelter and ramp strategies within a few rounds and to improve across four additional worlds.
Replay accuracy and cost were measured under identical model and hardware settings. Block-by-block SDPA replay is bitwise identical to the cached rollout, while full-sequence FlexAttention replay introduces a small relative error. The block-by-block approach also slightly reduces forward time with almost no change in peak allocated memory. Block-by-block SDPA replay exactly matches the cached rollout. Full-sequence FlexAttention replay adds a 3.99% relative L2 error. Replacing FlexAttention with block-by-block SDPA lowers forward time by 5.4% and changes peak memory by only 0.2%.
At steady state with a full cache, the transformer stage dominates inference cost for a served block of sixteen frames, while condition encoding and decoding add only a small overhead. The renderer sustains over 35 frames per second on one NVIDIA H100 GPU, with the tiny decoder completing in under 10 ms. The transformer stage, covering four denoising steps and cache publication, accounts for nearly all served-block latency. Condition encoding and tiny decoder with host transfer are minor contributors to total wall time. Steady-state throughput reaches over 35 frames per second for 480x832 video on a single H100.
The experiment tracks two physical tool-use milestones in hide-and-seek: shelter construction and using ramps to enter shelters. Pretrained visual agents reached both milestones within a handful of rounds while perceiving only through a real-time neural renderer and updating written playbooks, whereas self-play reinforcement learning from scratch required millions of training episodes. The comparison highlights rapid grounding of general commonsense priors into closed-loop physical execution rather than a direct sample-efficiency ratio. Shelter construction and ramp use emerged after only a few rounds for pretrained agents, while self-play reinforcement learning required tens to hundreds of millions of episodes. Pretrained agents achieved these tool-use milestones from first-person neural-rendered observations alone, without object coordinates or privileged world state, and refined tactics such as repositioning a ramp closer to a wall after an initial failed attempt.
Most worlds showed clear improvement across rounds. Companion dog engagement rose and plateaued at its highest observed level, bridge crossing time fell substantially, herding moved from three sheep penned to all four in every later round, and the quarry loader recovered from an early scoreless round to complete most of the task in the final round. Companion dog engagement improved from the initial round and ended at the highest recorded score. Bridge agents completed every round and cut their crossing time substantially across rounds. The herding pair penned all four sheep in every later round after falling short only in the first round. Quarry loader performance was uneven early but improved sharply in the final round, clearing rocks, delivering one, and parking within the time limit.
The experiments evaluate an interactive video generation and embodied agent pipeline across several levels. A block-by-block SDPA replay scheme matches cached rollouts exactly while slightly reducing forward time and leaving memory nearly unchanged, unlike full-sequence FlexAttention replay, which introduces numerical error. At steady state, the renderer sustains real-time video generation on a single H100 with the transformer stage dominating cost, while condition encoding and decoding remain minor. Pretrained visual agents reach physical tool-use milestones in only a few rounds from rendered first-person observations, in contrast to self-play reinforcement learning requiring massive training. Across multiple simulated worlds, agents generally improve over rounds, including more consistent herding, faster bridge crossing, and better quarry loading.