Command Palette
Search for a command to run...
プロアクティブエージェントの基礎:原理、技術レイヤ、およびPROACTIVITY-GYM
プロアクティブエージェントの基礎:原理、技術レイヤ、およびPROACTIVITY-GYM
Jio Oh Seunghyun Do Young-Jun Lee Steven Euijong Whang Dongyeop Kang
概要
プロアクティブLLMエージェントは、ユーザが求める前から、アイドル状態の計算資源を有益な支援に変えることができる。しかし、たとえ成果が正しくても、ユーザの文脈を読み誤ったり、確認コストを課したり、信頼を損なったりする可能性がある。本研究は、3つの統合的原則(3T)を中心に、プロアクティブLLMエージェントの設計・実現・評価のための基礎を提案する。すなわち、タスク能力(Task Capability、関連するニーズを先回りして有用な作業を正しく遂行する)、時間配分(Temporal Allocation、リソースの可用性と結果が必要となる時期に応じて計算資源を割り当てる)、信頼(Trust、ユーザの確信とエージェントへの適切な依存を維持する)である。我々はこれらの目的を、タスク範囲、先読み期間(anticipation horizon)、起動トリガー、処理タイミング、介入深度という5つの次元で構成される設計空間に結び付け、その選択を支えるために必要な状況およびシステムのモデル化(ユーザ表現と環境表現、バックボーンLLM、エージェントハーネスを含む)を明確化する。最後に、複数日にわたるシナリオ、状態を保持する環境、ペルソナ条件付きシミュレートユーザを含むシミュレーションベースの評価テストベッドであるPROACTIVITY-GYMを提案する。これにより、インタラクション全体にわたるプロアクティブ支援の帰結を評価できる。23種類のモデル・ハーネス構成にわたる評価では、3Tにおいて大きな性能差が明らかになり、LLM判定器がタスク能力と信頼をしばしば混同することが示された。30名の参加者による人間研究は、3Tの統合的最適化の重要性を示している。参加者は、成果が正しい場合でも介入の不整合後に信頼が急激に低下し、進行中の集中を保つために、たとえ不完全であっても睡眠時間帯の支援を好む。これらの知見は総合して、有用な作業、計算資源の配分、変化するユーザの信頼を統合的に考慮することによってプロアクティブエージェントを設計・評価することを支持する。
One-sentence Summary
Researchers from KAIST and the University of Minnesota propose foundations for designing, realizing, and evaluating proactive LLM agents around the joint 3T principles of Task Capability, Temporal Allocation, and Trust, and they introduce PROACTIVITY-GYM, a simulation-based testbed whose evaluations across 23 model-harness configurations and a 30-participant human study reveal sharp trust declines after intervention misalignment and a preference for sleep-time assistance even when imperfect.
Key Contributions
- Introduces a blueprint for proactive LLM agents built on three joint objectives: Task Capability, Temporal Allocation, and Trust (3T), which are often conflated or overlooked in existing work.
- Formalizes the proactivity design space along five dimensions: task scope, anticipation horizon, activation trigger, processing timing, and intervention depth, and specifies the situation and system modeling needed to realize these choices, including user and environment representations, backbone LLMs, and agent harnesses.
- Presents PROACTIVITY-GYM, a simulation-based evaluation testbed with multi-day scenarios, stateful environments, and persona-conditioned simulated users; evaluations across 23 model-harness configurations and a 30-participant human study reveal substantial performance gaps across the 3T objectives and show that task capability alone is insufficient for sustaining appropriate trust.
Introduction
Open-source agent frameworks and open-weight models now make it practical to run personal AI agents on consumer hardware, creating an opportunity for proactive assistance that uses idle compute before the user asks. However, current agent research and deployment remain mostly reactive, and existing evaluations often collapse failures in timing, intervention depth, and user trust into simple task success, which hides why proactive help may be ignored or rejected. The authors address this gap by proposing three joint design principles, Task Capability, Temporal Allocation, and Trust (3T), and by introducing PROACTIVITY-GYM, a multi-day simulation testbed with persona-conditioned feedback; they evaluate 23 model-harness configurations and conduct a 30-participant study showing that task performance alone is insufficient.
Dataset
The authors introduce PROACTIVITY-GYM as a synthetic, dynamic evaluation environment for proactive assistance. The data is purpose-built rather than collected from an existing corpus.
- Composition and scale: It contains 10 multi-day scenarios. Each scenario spans 7 to 10 simulated days and covers professional or everyday domains such as research, shopping, and business.
- Scenario schema: Each scenario includes a user goal, timestamped events and tasks, relevant tools organized into multiple sessions, latent user needs, and competing tasks with resource, availability, and deadline constraints.
- Persona variants: For each scenario, the authors construct three user personas that differ in preferred intervention depth. This yields 30 persona-scenario configurations.
- User simulator: A persona-conditioned user simulator provides implicit and explicit feedback, which can change later interactions and should be used to infer trust state.
- Dynamic and NOOP cases: Scenarios are stateful and action-dependent. Follow-up events depend on earlier user and agent actions and the resulting state. The authors also include NOOP cases where a proactive opportunity is irrelevant or already resolved and no intervention is needed.
- Processing and filtering: No filtering is described because the scenarios are manually constructed for evaluation. The main processing design is the addition of dynamic constraints and persona-conditioned feedback.
- Usage: The authors use this environment to evaluate proactive assistance under the 3T objectives, combining a simulated clock, stateful tool environments, and scheduled events. It tests how proactive actions affect later work, task progress, resource availability, and user state over time.
Method
The authors define proactivity as an agent's anticipatory behavior intended to address and fill the gaps of user needs without an explicit request. To ensure useful assistance, they introduce three design objectives (3T): Task Capability (TC), Temporal Allocation (TA), and Trust (TR). Task Capability is the ability to anticipate relevant user needs and correctly perform useful work. Temporal Allocation is the ability to allocate compute over time according to resource availability and when results are needed, distinguishing between interaction time and sleep time. Trust is the extent to which a user is confident in and willing to rely on an agent's recommendations, often proxied by intervention depth.
As shown in the figure below, these objectives jointly shape assistance in a workflow, such as deferring compute-heavy tasks to sleep time to avoid disrupting the user.
To guide agent design and evaluation, the authors combine these objectives into a weighted formulation:
A∈AmaxλTCSTC(A)+λTASTA(A)+λTRSTR(A)where A is the set of agent designs, S scores the design on each objective, and λ are weights summing to 1.
The authors organize proactive assistance into a design space spanning three connected decisions: what work to pursue, when to initiate and process it, and how far to proceed. Refer to the framework diagram for the five dimensions.
Identifying useful work involves choosing the task scope (within-task or out-of-task) and anticipation horizon (immediate, within-interaction, or out-of-interaction). Initiating and scheduling work depends on the activation trigger (user, event, or agent) and processing timing (interaction time or sleep time). Choosing intervention depth determines autonomy levels: prepare, suggest, or execute, reflecting estimated trust.
To realize such agents, the authors propose a system architecture connecting context, decisions, and feedback through situation and system modeling.
Situation modeling integrates interaction history, memory, and external observations to represent the current user and environment state, tracking goals, attention, preferences, trust, and resource availability. System modeling specifies how the backbone LLM and harness support proactive work. The harness orchestrates model calls, tools, memory, and permissions, tracking pending tasks and allocating compute.
At decision step t, the components select an action from the available context ct:
zt=R(ct),at∼πA(⋅∣zt,ct),ct+1=U(ct,at,ot+1)Here, R builds the situation representation zt, πA selects an action, and U updates the context with the action trace and new observations ot+1, including user feedback.
Experiment
The experiments evaluate nine models across OpenClaw, Claude Code, and Codex harnesses using task completion, temporal allocation, intervention-depth alignment, and trust ratings, and include a 30-person human study with scenario reviews and simulated week-long logs. Benchmark results show that even the strongest agents leave substantial gaps in temporal allocation and intervention alignment, and that task completion gains do not consistently improve those dimensions; LLM-based trust ratings also appear tolerant of misaligned interventions. The human study reinforces that people value correctness but strongly prefer intervention alignment, that deferring assistance to periods like sleep time can make proactive help more acceptable, and that trust losses from misalignment are larger and harder to recover than gains.
The example scenario contrasts an agent that asks for confirmation before adding a language review session and reminder with an agent that performs the same correct actions without approval. Participants valued task correctness highly, but they also valued alignment with the user's standing request, and trust was sensitive to intervention misalignment. Proactive assistance was more acceptable when scheduled for low-disruption periods such as sleep time, even if it required later correction. When content quality was held constant, participants preferred agents that respected the user's approval preference. Misaligned intervention behavior reduced trust even when the agent's task outcomes remained correct. Trust declined with repeated misalignment and recovered only partially after a return to aligned behavior. Participants often preferred sleep-time assistance over immediate assistance, even when the immediate action was correct and required little revision.
This experiment compared agents that requested confirmation before adding a language review session and reminder against agents that acted identically but without approval. It examined how proactive assistance, timing, and alignment with user preferences affect trust and acceptability. Participants valued both task correctness and respect for a standing request; trust dropped with repeated misaligned interventions even when outcomes were correct, and recovered only partially after realignment. They often preferred low-disruption sleep-time assistance over immediate correct actions that required little revision.