HyperAIHyperAI

Command Palette

Search for a command to run...

痕跡からエージェンティックな世界へ:対話型環境シミュレーションのためのエージェンティック言語世界モデル

Quanyu Long Xiao Chen Jianda Chen Haozhen Zhang Qisheng Hu Jianzhu Bao Wenya Wang

概要

現実的な環境レプリカはLLMエージェントの訓練と評価においてますます有用になっているが、元のシステムにはアクセスできない場合や再現が現実的でない場合がある。本研究ではエージェンティック言語世界モデリングを検討する。これは実行可能な環境を再構築するのではなく、世界モデルエージェントがタスクエージェントにとっての環境として機能し、忠実で状態を持つシミュレーションを支えるというものである。我々はこのパラダイムを、元のシステムは利用できないが過去の対話痕跡にはアクセスできる設定向けの学習不要フレームワークTrace2Envとして具体化する。Trace2Envはこれらの痕跡を、環境スキーマ、根拠に基づく証拠、帰納された行動知識を含む再利用可能な環境ワールドブックへ再構成する。実行時には、世界モデルエージェントが永続的なエピソード状態とともにワールドブックを能動的に参照し、各行動の観測と持続的な状態効果を推論する。9つの環境にわたり、Trace2Envは従来のプロンプトベース言語世界モデルと比較して、次観測の忠実度と長期的対話の一貫性の両方を改善する。複数ターンの対話では、Trace2Envに対して生成されたタスクエージェントの行動は、実環境で再実行した場合により高い頻度で有効であり続ける。これは、そのシミュレーション動態が連続するターンにわたって以前の行動の結果をよりよく保持することを示している。これらの結果は、元の実行可能システムを再構築することなく現実的な環境レプリカを構築するための代替的方向性として、エージェンティック言語世界モデリングを確立するものである。

One-sentence Summary

Researchers from Nanyang Technological University and The Hong Kong Polytechnic University propose Trace2Env, a learning-free agentic language world modeling framework that reconstructs historical interaction traces into a reusable environment worldbook, allowing a world model agent to consult environment schemas, grounded evidence, and induced behavioral knowledge alongside persistent episodic state, thereby improving next-observation fidelity and long-horizon interaction consistency across nine environments compared with conventional prompt-based language world models.

Key Contributions

  • The paper introduces agentic language world modeling, a paradigm where a world model agent serves as a stateful environment simulator for a task agent instead of requiring an executable replica of the original system.
  • The paper presents Trace2Env, a learning-free framework that reconstructs historical interaction traces into a reusable environment worldbook containing schemas, grounded evidence, and induced behavioral knowledge, then uses that worldbook with persistent episodic state to infer observations and lasting state effects.
  • Across nine environments, Trace2Env improves next-observation fidelity and long-horizon interaction consistency over prompt-based language world models, and task-agent actions generated against Trace2Env remain valid more often when replayed in the real environment.

Introduction

Interactive agents require environments that respond to actions and preserve their consequences, but building realistic replicas of terminals, software workspaces, or enterprise systems often requires code, data, or infrastructure that is unavailable. Historical interaction traces are often still accessible and contain behavioral evidence, yet prior trace-based approaches either produce skills for the acting agent or reconstruct executable workspaces, which may still depend on missing implementation details. Prompt-based language world models also flatten environment knowledge and long-horizon state into a single context, making them unreliable for maintaining state continuity over many turns. The authors address this by proposing Trace2Env, a learning-free framework that turns historical traces into a structured, non-executable environment worldbook and uses an agentic language world model to actively inspect that knowledge, maintain persistent episode state and memory, and simulate environment responses.

Method

The authors introduce Trace2Env, a framework designed for trace-based environment reconstruction that transfers behavioral knowledge across episodes without importing episode-specific facts. The system operates in two distinct phases: an offline phase that reconstructs an environment worldbook from recorded traces, and an online phase that combines this fixed knowledge with mutable episode state and interaction memory. A world model agent proposes each transition during the online rollout, while a shared harness controls which state changes are committed.

As shown in the figure below, the overall pipeline is divided into environment reconstruction, worldbook storage, and agentic simulation.

During the offline phase, the authors leverage recorded transitions to build a reusable environment worldbook. A recorded transition provides concrete evidence about a single interaction, but simulation requires knowledge that generalizes to unseen situations. The offline constructor processes the build traces to produce the worldbook KE=Constructϕ(Dbuild)K_{\mathcal{E}} = \mathrm{Construct}_{\phi}(\mathcal{D}_{\text{build}})KE​=Constructϕ​(Dbuild​). This worldbook comprises four complementary components. First, schemas describe the action interface and the state variables the simulator can maintain. Second, grounded evidence retains recorded transitions, selected demonstrations, and their original observations. The constructor aligns each action with its result to extract observed facts, candidate state effects, outcomes, and uncertainty, while withholding effects with ambiguous attribution. Third, induced abstractions capture recurring behaviors across traces, including conditional effects, state constraints, observation contracts, and descriptive conventions. The constructor proposes behavioral rules and reviews them against supporting and contrasting cases to refine these abstractions. Finally, provenance records link each artifact back to the specific traces and turns that support it, allowing the simulator to inspect the environment at varying levels of specificity.

To distinguish knowledge that transfers across episodes from facts established only within the current interaction, Trace2Env maintains a persistent episode workspace. Treating all signals as a single flat history would make their scope ambiguous. Instead, the workspace is defined as Wt=(KE,s^t,Mt)\mathcal{W}_{t} = (K_{\mathcal{E}}, \hat{s}_{t}, M_{t})Wt​=(KE​,s^t​,Mt​), where KEK_{\mathcal{E}}KE​ is the immutable environment worldbook shared across episodes, s^t\hat{s}_{t}s^t​ is the mutable represented state of the current episode, and MtM_{t}Mt​ is the episodic interaction memory. This separation ensures that cross-episode evidence informs how the environment behaves without establishing what exists in the current episode, treating missing state entries as unknown rather than absent.

In the online agentic simulation phase, the system produces the next transition in two stages. Given an action ata_tat​, the world model agent starts with a compact view of the current episode and selectively inspects additional information from the workspace. Because cross-episode evidence introduces the challenge that relevance does not imply applicability, the system separates retrieval from applicability. For each retrieved worldbook entry eee, an applicability gate assigns a classification:

gtapp(e)=Gapp(e∣at,Wt)∈{supporting, uncertain, format-only}g_{t}^{\mathrm{app}}(e) = G_{\mathrm{app}}(e \mid a_{t}, \mathcal{W}_{t}) \in \{\text{supporting, uncertain, format-only}\}gtapp​(e)=Gapp​(e∣at​,Wt​)∈{supporting, uncertain, format-only}

Supporting evidence informs concrete behavior, uncertain evidence is used cautiously, and format-only evidence contributes response structure without establishing episode-specific facts. The agent alternates between inspecting the workspace and reasoning about the transition until it is ready to externalize its decision as a transition proposal qt=(Δt,o~t+1)∼Pθ(⋅∣dE,at,Wt)q_{t} = (\Delta_{t}, \tilde{o}_{t+1}) \sim P_{\theta}(\cdot \mid d_{\mathcal{E}}, a_{t}, \mathcal{W}_{t})qt​=(Δt​,o~t+1​)∼Pθ​(⋅∣dE​,at​,Wt​). Here, Δt\Delta_{t}Δt​ is an ordered list of candidate state effects, and o~t+1\tilde{o}_{t+1}o~t+1​ is the proposed observation.

Rather than modifying the episode directly, the proposal is passed to the shared harness for validation. The harness applies a validation gate:

gtval(qt)=Gval(qt∣at,Wt)∈{accept, reject}g_{t}^{\mathrm{val}}(q_{t}) = G_{\mathrm{val}}(q_{t} \mid a_{t}, \mathcal{W}_{t}) \in \{\text{accept, reject}\}gtval​(qt​)=Gval​(qt​∣at​,Wt​)∈{accept, reject}

This gate checks the candidate effects on a state copy against the schemas, rule support, and applicable state constraints. Only accepted proposals are committed. The harness commits the staged effects and returns the resolved observation (s^t+1,o^t+1)=H(Wt,at,qt)(\hat{s}_{t+1}, \hat{o}_{t+1}) = \mathcal{H}(\mathcal{W}_{t}, a_{t}, q_{t})(s^t+1​,o^t+1​)=H(Wt​,at​,qt​). Once committed, the updated state and interaction memory persist into the next turn, providing continuity across the online rollout.

Experiment

The experiments evaluate Trace2Env across nine environments in two settings: next-observation prediction on AgentWorldBench and EnvScaler, and multi-turn interaction on ALFWorld and SciWorld. Comparisons with direct prompting, trace retrieval, worldbook prompting, and harness-only baselines show that the full system improves prediction quality and consistency, with reconstructed worldbook knowledge and the runtime contributing complementary gains. Long-horizon interaction results indicate that high simulated success alone does not guarantee a faithful world model, while Trace2Env achieves stronger transfer back to real environments by preserving grounded state constraints. Ablations further show that worldbook evidence and abstraction are complementary and that fidelity improves with more construction traces, and qualitative results highlight how persistent state tracking prevents incorrect transitions from derailing interaction.

For the gpt-5.6-sol backbone, Trace2Env achieves the highest average next-observation prediction score across the seven evaluated environments. Prompting with retrieved traces or worldbook context also improves over direct prompting, while the harness-only variant shows smaller and less consistent gains. The strongest scores occur in application-style environments such as Food, Shopping, and Benefits. Trace2Env leads overall and records the best score in most environments, with especially large gains over direct prompting in Terminal and Benefits. Worldbook prompting and trace RAG prompting both outperform direct prompting, with worldbook prompting slightly ahead on average and in most environments. Harness-only results are mixed, improving over direct prompting in some environments but trailing in others.

Direct Prompting achieves high task success in simulated world models, but its action sequences often fail when replayed in the real environment. Trace2Env has lower simulated success but much higher world-model-to-real consistency. The results indicate that simulated success alone is not a reliable measure of world model fidelity. In ALFWorld, Direct Prompting drops from 97% simulated success to 3% when its induced actions are replayed in the real environment. Trace2Env maintains substantially higher W2R consistency in both ALFWorld and SciWorld even though its simulated success is lower than Direct Prompting.

On Terminal, evidence provides the stronger standalone contribution to worldbook performance, while abstraction is most useful when paired with evidence rather than used alone. The full combination of abstraction and evidence achieves the best result. Worldbook performance also improves as more construction traces are used, though construction cost rises with trace volume, so a moderate trace count is chosen as a practical trade-off. Evidence improves the schema-only variant substantially, whereas abstraction alone does not help and slightly lowers performance. Adding evidence to an abstracted worldbook produces a larger gain than adding abstraction on top of evidence, and the full worldbook performs best. Performance becomes more stable and continues rising with additional construction traces, but one-time construction cost increases with trace volume.

The experiments evaluate trace-based environment induction against direct prompting and retrieval or worldbook baselines for next-observation prediction across seven environments, finding that Trace2Env performs best overall and especially in application-style settings such as Food, Shopping, and Benefits. A second study on simulated versus real replay shows that direct prompting can achieve high simulated success but poor world-model-to-real consistency, while Trace2Env trades some simulated success for substantially more faithful action transfer in ALFWorld and SciWorld. Ablations on Terminal indicate that evidence contributes more than abstraction to worldbook performance, the full combination of abstraction and evidence works best, and using more construction traces improves stability at added construction cost.


AIでAIを構築

アイデアからローンチまで — 無料のAIコーディング支援、すぐに使える環境、最高のGPU価格でAI開発を加速。

AI コーディング補助
すぐに使える GPU
最適な料金体系

HyperAI Newsletters

最新情報を購読する
北京時間 毎週月曜日の午前9時 に、その週の最新情報をメールでお届けします
メール配信サービスは MailChimp によって提供されています