Command Palette
Search for a command to run...
أسس الوكلاء الاستباقيين: المبادئ والطبقات التقنية وPROACTIVITY-GYM
أسس الوكلاء الاستباقيين: المبادئ والطبقات التقنية وPROACTIVITY-GYM
Jio Oh Seunghyun Do Young-Jun Lee Steven Euijong Whang Dongyeop Kang
الملخص
يمكن للوكلاء الاستباقيين القائمين على نماذج اللغة الكبيرة (LLM) أن يحوّلوا القدرة الحاسوبية الخاملة إلى دعم مفيد قبل أن يطلب المستخدم ذلك. ومع ذلك، حتى العمل الصحيح قد يُساء فهم سياق المستخدم، أو يفرض تكاليف مراجعة، أو يُضعف الثقة. يقترح هذا العمل أسسًا لتصميم الوكلاء الاستباقيين القائمين على LLM وتحقيقهم وتقييمهم حول ثلاثة مبادئ مشتركة (3T): قدرة المهمة (Task Capability)، أي استباق الاحتياجات ذات الصلة وأداء عمل مفيد على نحو صحيح؛ والتوزيع الزمني (Temporal Allocation)، أي تخصيص القدرة الحاسوبية وفقًا لتوافر الموارد ووقت الحاجة إلى النتائج؛ والثقة (Trust)، أي الحفاظ على ثقة المستخدمين واعتمادهم المناسب على الوكيل. ونربط هذه الأهداف بفضاء تصميمي ينتظم حول خمسة أبعاد هي: نطاق المهمة، وأفق الاستباق، ومحفز التفعيل، وتوقيت المعالجة، وعمق التدخل؛ ونحدد نمذجة الموقف والنظام اللازمة لدعم خياراته، بما في ذلك تمثيلات المستخدم والبيئة، ونماذج LLM الأساسية، وهياكل تشغيل الوكيل. وأخيرًا، نقترح PROACTIVITY-GYM، وهي بيئة تقييم قائمة على المحاكاة تتضمن سيناريوهات متعددة الأيام، وبيئات ذات حالة، ومستخدمين محاكَين مشروطين بأنماط شخصيات، ويمكنها تقييم عواقب المساعدة الاستباقية عبر التفاعلات. تكشف التقييمات عبر 23 تكوينًا من النماذج وهياكل التشغيل عن فجوات كبيرة في الأداء عبر المبادئ الثلاثة (3T)، وتبيّن أن المحكّمات القائمة على نماذج LLM غالبًا ما تخلط بين قدرة المهمة والثقة. وتُظهر دراسة بشرية شملت 30 مشاركًا أهمية التحسين المشترك للمبادئ الثلاثة (3T): إذ يُظهر المشاركون انخفاضًا حادًا في الثقة بعد عدم توافق التدخل رغم صحة النتائج، ويفضلون المساعدة أثناء النوم، حتى عندما تكون غير مثالية، للحفاظ على التركيز المستمر. وتدعم هذه النتائج مجتمعةً تصميم الوكلاء الاستباقيين وتقييمهم عبر النظر المشترك إلى العمل المفيد، وتخصيص الحوسبة، وتطور ثقة المستخدم.
One-sentence Summary
Researchers from KAIST and the University of Minnesota propose foundations for designing, realizing, and evaluating proactive LLM agents around the joint 3T principles of Task Capability, Temporal Allocation, and Trust, and they introduce PROACTIVITY-GYM, a simulation-based testbed whose evaluations across 23 model-harness configurations and a 30-participant human study reveal sharp trust declines after intervention misalignment and a preference for sleep-time assistance even when imperfect.
Key Contributions
- Introduces a blueprint for proactive LLM agents built on three joint objectives: Task Capability, Temporal Allocation, and Trust (3T), which are often conflated or overlooked in existing work.
- Formalizes the proactivity design space along five dimensions: task scope, anticipation horizon, activation trigger, processing timing, and intervention depth, and specifies the situation and system modeling needed to realize these choices, including user and environment representations, backbone LLMs, and agent harnesses.
- Presents PROACTIVITY-GYM, a simulation-based evaluation testbed with multi-day scenarios, stateful environments, and persona-conditioned simulated users; evaluations across 23 model-harness configurations and a 30-participant human study reveal substantial performance gaps across the 3T objectives and show that task capability alone is insufficient for sustaining appropriate trust.
Introduction
Open-source agent frameworks and open-weight models now make it practical to run personal AI agents on consumer hardware, creating an opportunity for proactive assistance that uses idle compute before the user asks. However, current agent research and deployment remain mostly reactive, and existing evaluations often collapse failures in timing, intervention depth, and user trust into simple task success, which hides why proactive help may be ignored or rejected. The authors address this gap by proposing three joint design principles, Task Capability, Temporal Allocation, and Trust (3T), and by introducing PROACTIVITY-GYM, a multi-day simulation testbed with persona-conditioned feedback; they evaluate 23 model-harness configurations and conduct a 30-participant study showing that task performance alone is insufficient.
Dataset
The authors introduce PROACTIVITY-GYM as a synthetic, dynamic evaluation environment for proactive assistance. The data is purpose-built rather than collected from an existing corpus.
- Composition and scale: It contains 10 multi-day scenarios. Each scenario spans 7 to 10 simulated days and covers professional or everyday domains such as research, shopping, and business.
- Scenario schema: Each scenario includes a user goal, timestamped events and tasks, relevant tools organized into multiple sessions, latent user needs, and competing tasks with resource, availability, and deadline constraints.
- Persona variants: For each scenario, the authors construct three user personas that differ in preferred intervention depth. This yields 30 persona-scenario configurations.
- User simulator: A persona-conditioned user simulator provides implicit and explicit feedback, which can change later interactions and should be used to infer trust state.
- Dynamic and NOOP cases: Scenarios are stateful and action-dependent. Follow-up events depend on earlier user and agent actions and the resulting state. The authors also include NOOP cases where a proactive opportunity is irrelevant or already resolved and no intervention is needed.
- Processing and filtering: No filtering is described because the scenarios are manually constructed for evaluation. The main processing design is the addition of dynamic constraints and persona-conditioned feedback.
- Usage: The authors use this environment to evaluate proactive assistance under the 3T objectives, combining a simulated clock, stateful tool environments, and scheduled events. It tests how proactive actions affect later work, task progress, resource availability, and user state over time.
Method
The authors define proactivity as an agent's anticipatory behavior intended to address and fill the gaps of user needs without an explicit request. To ensure useful assistance, they introduce three design objectives (3T): Task Capability (TC), Temporal Allocation (TA), and Trust (TR). Task Capability is the ability to anticipate relevant user needs and correctly perform useful work. Temporal Allocation is the ability to allocate compute over time according to resource availability and when results are needed, distinguishing between interaction time and sleep time. Trust is the extent to which a user is confident in and willing to rely on an agent's recommendations, often proxied by intervention depth.
As shown in the figure below, these objectives jointly shape assistance in a workflow, such as deferring compute-heavy tasks to sleep time to avoid disrupting the user.
To guide agent design and evaluation, the authors combine these objectives into a weighted formulation:
A∈AmaxλTCSTC(A)+λTASTA(A)+λTRSTR(A)where A is the set of agent designs, S scores the design on each objective, and λ are weights summing to 1.
The authors organize proactive assistance into a design space spanning three connected decisions: what work to pursue, when to initiate and process it, and how far to proceed. Refer to the framework diagram for the five dimensions.
Identifying useful work involves choosing the task scope (within-task or out-of-task) and anticipation horizon (immediate, within-interaction, or out-of-interaction). Initiating and scheduling work depends on the activation trigger (user, event, or agent) and processing timing (interaction time or sleep time). Choosing intervention depth determines autonomy levels: prepare, suggest, or execute, reflecting estimated trust.
To realize such agents, the authors propose a system architecture connecting context, decisions, and feedback through situation and system modeling.
Situation modeling integrates interaction history, memory, and external observations to represent the current user and environment state, tracking goals, attention, preferences, trust, and resource availability. System modeling specifies how the backbone LLM and harness support proactive work. The harness orchestrates model calls, tools, memory, and permissions, tracking pending tasks and allocating compute.
At decision step t, the components select an action from the available context ct:
zt=R(ct),at∼πA(⋅∣zt,ct),ct+1=U(ct,at,ot+1)Here, R builds the situation representation zt, πA selects an action, and U updates the context with the action trace and new observations ot+1, including user feedback.
Experiment
The experiments evaluate nine models across OpenClaw, Claude Code, and Codex harnesses using task completion, temporal allocation, intervention-depth alignment, and trust ratings, and include a 30-person human study with scenario reviews and simulated week-long logs. Benchmark results show that even the strongest agents leave substantial gaps in temporal allocation and intervention alignment, and that task completion gains do not consistently improve those dimensions; LLM-based trust ratings also appear tolerant of misaligned interventions. The human study reinforces that people value correctness but strongly prefer intervention alignment, that deferring assistance to periods like sleep time can make proactive help more acceptable, and that trust losses from misalignment are larger and harder to recover than gains.
The example scenario contrasts an agent that asks for confirmation before adding a language review session and reminder with an agent that performs the same correct actions without approval. Participants valued task correctness highly, but they also valued alignment with the user's standing request, and trust was sensitive to intervention misalignment. Proactive assistance was more acceptable when scheduled for low-disruption periods such as sleep time, even if it required later correction. When content quality was held constant, participants preferred agents that respected the user's approval preference. Misaligned intervention behavior reduced trust even when the agent's task outcomes remained correct. Trust declined with repeated misalignment and recovered only partially after a return to aligned behavior. Participants often preferred sleep-time assistance over immediate assistance, even when the immediate action was correct and required little revision.
This experiment compared agents that requested confirmation before adding a language review session and reminder against agents that acted identically but without approval. It examined how proactive assistance, timing, and alignment with user preferences affect trust and acceptability. Participants valued both task correctness and respect for a standing request; trust dropped with repeated misaligned interventions even when outcomes were correct, and recovered only partially after realignment. They often preferred low-disruption sleep-time assistance over immediate correct actions that required little revision.