Command Palette
Search for a command to run...
AgentGarten: Code-Welten für sich weiterentwickelnde Agenten
AgentGarten: Code-Welten für sich weiterentwickelnde Agenten
Zusammenfassung
Interaktive virtuelle Welten ermöglichen es Agenten, durch Erkundung und Interaktion zu lernen. Was Agenten lernen können, wird durch die Umgebungen begrenzt, in denen sie üben; diese müssen getreu sein – mit konsistentem Zustand, konsistenten Regeln und konsistenter Dynamik – und realistisch, mit Beobachtungen, die den visuellen Verteilungen der realen Welt folgen. Beides über unterschiedlichste Welten hinweg zu erreichen, bleibt ein Engpass. Wir stellen AgentGarten vor, ein Framework, das Simulatoren und Game-Engines mit einem gemeinsamen neuronalen Renderer koppelt, um interaktive Echtzeitumgebungen zu erstellen. Die Simulations-Backends verwalten einen persistenten Weltzustand und führen programmdefinierte Interaktionsregeln aus, während der Renderer visuelle Beobachtungen aus strukturierten Bedingungen erzeugt, die über eine gemeinsame Schnittstelle exportiert werden. Für den Aufbau des neuronalen Renderers adaptieren wir ein vortrainiertes Videomodell an Geometriebedingungen, destillieren es mit unserem vorgeschlagenen Adversarial Forcing und optimieren die Inferenz für Echtzeitinteraktion. Adversarial Forcing macht das Vorbefüllen des Verlaufs durch exaktes Replay differenzierbar, sodass Verluste aus späteren Vorhersagen die Art und Weise aktualisieren, wie der Renderer frühere Beobachtungen kodiert, und fügt eine adversarielle Überwachung mit realen Daten hinzu, um die visuelle Qualität zu verbessern. In AgentGarten nehmen Agenten die Welt über visuelle Beobachtungen wahr, interagieren mit ihr in Echtzeit und verbessern sich, indem sie jede Erfahrungsrunde zu Playbooks destillieren, die nachfolgende Agenten übernehmen und verfeinern. Unsere empirische Studie zeigt einen erheblichen Gewinn an Lerneffizienz: Agenten lernen bereits aus nur 4 Runden, verglichen mit Millionen Runden bei einem herkömmlichen Reinforcement-Learning-Ansatz. Da neue Welten als Code geschrieben und über dieselbe Schnittstelle gerendert werden können, lassen sich Umgebungen sowohl hinsichtlich ihrer Anzahl als auch ihrer Schwierigkeit gemeinsam mit ihren Agenten skalieren – ein Schritt hin zu Agenten, die sich durch interaktive Erfahrung kontinuierlich weiterentwickeln.
One-sentence Summary
Researchers at MirroS, Tsinghua University, and Peking University introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments, adapts a pretrained video model using Adversarial Forcing and real-data adversarial supervision, and enables agents to learn from just four rounds compared with millions for a conventional reinforcement learning counterpart.
Key Contributions
- AgentGarten couples simulator or game-engine backends with a shared neural renderer to build real-time interactive environments; the backends maintain persistent world state and program-defined interaction rules, while the renderer converts structured conditions exported through a common interface into visual observations.
- Adversarial Forcing adapts a pretrained video model to geometry conditions, makes history prefilling differentiable through exact replay so later prediction losses update how prior observations are encoded, and adds real-data adversarial supervision to improve visual quality.
- The empirical study demonstrates that agents distill each round of experience into playbooks that later agents inherit and refine, learning from just 4 rounds compared with millions for a conventional reinforcement learning counterpart; in hide-and-seek, shelters appeared by round 4 and ramp crossings by round 10, and the same procedure improved outcomes in four further worlds.
Introduction
The authors target the need for scalable training environments in which agents learn from long trajectories of actions and visual observations. Such environments must preserve the consequences of past actions while allowing variation in layouts, objects, and task rules. Prior simulators and game engines offer explicit, inspectable state and programmable rules, but expanding their visual diversity is costly because it requires new 3D assets and rendering pipelines. Conversely, video world models synthesize rich observations, but their task-relevant state and transition rules remain implicit and hard to inspect, edit, or test. The authors introduce AgentGarten, a framework that combines programmable scene programs with a shared real-time neural renderer. Each scene runs in a simulator or game engine that maintains persistent state and executes interaction rules, while the renderer converts structured engine conditions into visual observations. To make this renderer interactive over extended rollouts, they adapt a pretrained bidirectional video model to block-causal generation and distill it with Adversarial Forcing, using exact replay for history gradients and an exact adversarial regularization scheme to preserve long-rollout visual quality.
Dataset
The authors describe a dataset of executable “code worlds” rather than a fixed set of images or videos. Each code world pairs an executable scene program with an appearance reference x0, and the scene program stores only coarse geometry needed for layout, silhouettes, occlusion, and motion dynamics.
-
Sources and composition
- Image-sourced code worlds: A single input image is used as both the appearance reference x0 and the layout reference. Perception and generation models extract instance masks, monocular depth, camera intrinsics, 3D bounding boxes, and object meshes.
- Text-sourced code worlds: The coding agent writes a scene program directly from geometric primitives and engine assets. An image editing model converts the initial render into x0 while preserving scene layout.
- Environment library: The collection includes navigation, manipulation, tool use, and multi-agent interaction scenes.
- Interactive browser games: The authors add bowling, a penalty shootout, and a crate-vault puzzle. Player input via keyboard, mouse, or gamepad advances the code world, while the rendered stream is the only view.
-
Processing and refinement
- For image-based scenes, the coding agent fits extracted assets to estimated geometry, completes unobserved regions as plausible spatial extensions, and binds physical properties and interaction rules.
- For text-based scenes, the initial render is transformed into x0 with layout-preserving image editing.
- Both modalities use an interactive refinement loop: the agent evaluates rollouts from multiple probe cameras, inspects engine verification queries for contact, collision clearance, and occlusion, and repairs faulty poses, unsupported structures, or ambiguous motions.
-
Schema and usage
- Each code world consists of an executable scene program and an appearance reference x0.
- Different scene programs and compatible engines form code worlds with a shared renderer.
- Because state evolution is explicit, recorded rollouts can be replayed and rendered again, including re-shooting from a different camera or in a different visual style without changing what happened.
-
Statistics and splits
- The provided excerpt does not report exact subset sizes, source dataset counts, filtering thresholds, training/evaluation split ratios, mixture ratios, or cropping details.
Method
The authors formulate a code world as a combination of a scene program, an execution engine, and a neural renderer. The program specifies the scene and interaction rules. Let st denote the scene state, πt the camera pose, and at the actions. The engine updates the state and captures structured conditions ct:
(st+1,πt+1)=fp(st,πt,at),ct=hp(st,πt)Given an appearance reference x0 and text description y, the neural renderer generates observations xt from visual history x<t:
xt=Rθ(x<t,c≤t,x0,y)This separation ensures that rendering errors do not accumulate in the state, allowing recorded state trajectories to be rendered under different appearances or cameras without altering the underlying events.
As shown in the interaction diagram below:
To construct these environments, a coding agent builds a code world from a single image or text description. When building from an image, the input serves simultaneously as the appearance reference and the reference layout. Perception and generation models lift the visible content into instance masks, monocular depth, camera intrinsics, and object meshes. The agent writes a scene program, fits these assets to the estimated geometry, and refines the scene through an interactive feedback loop. It evaluates test rollouts from multiple probe cameras and iteratively repairs faulty poses, unsupported structures, or ambiguous motions before deployment.
The process of building a code world from a single image is illustrated here:
The authors train a geometry-conditioned autoregressive video model to implement the neural renderer. The backbone separates understanding and generation towers. The frozen understanding tower encodes the caption into per-layer keys and values consumed by the generation tower. Geometry conditions, such as depth and surface normals, are encoded and concatenated with RGB tokens along the sequence dimension for joint attention. This preserves geometric alignment while allowing cross-token interactions without imposing a pixel-aligned inductive bias.
The final training stage, Adversarial Forcing, distills the teacher-forced model into a few-step renderer trained on its own rollouts. It combines distribution matching against the bidirectional teacher, exact history-gradient replay, and real-data adversarial training. To allow later losses to update the history-encoding computation without the memory cost of a fully differentiable rollout, the authors introduce exact replay. This method decouples rollout and gradient propagation using two passes. A no-gradient rollout records each block's input and clean output, and a differentiable replay recomputes both history and predictions with the specific visibility pattern.
The block-level attention masks for exact replay are depicted below:
For adversarial training, a trainable head aggregates intermediate features from the frozen teacher backbone into a discriminator logit. To regularize the discriminator with R1 and R2 penalties without requiring double backward through fused attention kernels, the authors compute the exact penalty and its exact head-parameter gradient. They utilize vector-Jacobian products to obtain gradients and Jacobian-vector products to carry directions forward through the frozen backbone, avoiding the materialization of the full Jacobian matrix.
The computation flow for exact R1/R2 regularization is shown below:
For continuous inference within the interaction loop, the renderer utilizes a bounded key-value cache. The cache maintains a permanent sink prefix, which includes the appearance reference, and a sliding window of recent history. Older frames are evicted as new blocks are denoised and published to the cache. This structure supports rollouts exceeding the training horizon through top-aligned rotary position remapping, preserving relative temporal distances while geometry tracks the true simulation timeline.
The structure of the inference cache is visualized below:
Experiment
The experiments evaluate visual quality, replay accuracy, and inference throughput on one NVIDIA H100 GPU with BF16 computation, 480×832 outputs, four-frame latent blocks, and a four-step sampler. Adversarial Forcing maintains natural textures and fine detail in long rollouts while Self Forcing produces repetitive patterns and loses detail, and block-by-block SDPA replay is bitwise identical and slightly faster than cached rollouts while FlexAttention replay introduces small error. The renderer delivers real-time 480×832 video generation, and agent experiments demonstrate that closed-loop grounding from neural-rendered first-person observations enables hide-and-seek agents to develop shelter and ramp strategies within a few rounds and to improve across four additional worlds.
Replay accuracy and cost were measured under identical model and hardware settings. Block-by-block SDPA replay is bitwise identical to the cached rollout, while full-sequence FlexAttention replay introduces a small relative error. The block-by-block approach also slightly reduces forward time with almost no change in peak allocated memory. Block-by-block SDPA replay exactly matches the cached rollout. Full-sequence FlexAttention replay adds a 3.99% relative L2 error. Replacing FlexAttention with block-by-block SDPA lowers forward time by 5.4% and changes peak memory by only 0.2%.
At steady state with a full cache, the transformer stage dominates inference cost for a served block of sixteen frames, while condition encoding and decoding add only a small overhead. The renderer sustains over 35 frames per second on one NVIDIA H100 GPU, with the tiny decoder completing in under 10 ms. The transformer stage, covering four denoising steps and cache publication, accounts for nearly all served-block latency. Condition encoding and tiny decoder with host transfer are minor contributors to total wall time. Steady-state throughput reaches over 35 frames per second for 480x832 video on a single H100.
The experiment tracks two physical tool-use milestones in hide-and-seek: shelter construction and using ramps to enter shelters. Pretrained visual agents reached both milestones within a handful of rounds while perceiving only through a real-time neural renderer and updating written playbooks, whereas self-play reinforcement learning from scratch required millions of training episodes. The comparison highlights rapid grounding of general commonsense priors into closed-loop physical execution rather than a direct sample-efficiency ratio. Shelter construction and ramp use emerged after only a few rounds for pretrained agents, while self-play reinforcement learning required tens to hundreds of millions of episodes. Pretrained agents achieved these tool-use milestones from first-person neural-rendered observations alone, without object coordinates or privileged world state, and refined tactics such as repositioning a ramp closer to a wall after an initial failed attempt.
Most worlds showed clear improvement across rounds. Companion dog engagement rose and plateaued at its highest observed level, bridge crossing time fell substantially, herding moved from three sheep penned to all four in every later round, and the quarry loader recovered from an early scoreless round to complete most of the task in the final round. Companion dog engagement improved from the initial round and ended at the highest recorded score. Bridge agents completed every round and cut their crossing time substantially across rounds. The herding pair penned all four sheep in every later round after falling short only in the first round. Quarry loader performance was uneven early but improved sharply in the final round, clearing rocks, delivering one, and parking within the time limit.
The experiments evaluate an interactive video generation and embodied agent pipeline across several levels. A block-by-block SDPA replay scheme matches cached rollouts exactly while slightly reducing forward time and leaving memory nearly unchanged, unlike full-sequence FlexAttention replay, which introduces numerical error. At steady state, the renderer sustains real-time video generation on a single H100 with the transformer stage dominating cost, while condition encoding and decoding remain minor. Pretrained visual agents reach physical tool-use milestones in only a few rounds from rendered first-person observations, in contrast to self-play reinforcement learning requiring massive training. Across multiple simulated worlds, agents generally improve over rounds, including more consistent herding, faster bridge crossing, and better quarry loading.