Command Palette
Search for a command to run...
من الدليل إلى الفعل: كيف يفشل الوكلاء المستخدمون للأدوات
من الدليل إلى الفعل: كيف يفشل الوكلاء المستخدمون للأدوات
Hongzhan Lin Shidong Cao Ziyang Luo Wenhao Chai Mong-Li Lee Wynne Hsu
الملخص
يُحدث الوكلاء المستخدمون للأدوات تغييرات ذات عواقب في الحالة الخارجية، غير أن النتائج الصحيحة لا تضمن أن أفعالهم كانت مدعومة بأدلة مثبتة مسبقًا. ندرس مواضع انهيار سلسلة الانتقال من الدليل إلى الفعل أثناء انتقال الوكلاء من اتخاذ قرار الفعل من عدمه إلى تنفيذ إجراءات مفردة وسير عمل مترابطة. عبر عشرة تكوينات من النموذج ومنصة التنفيذ، يمكن أن يتعايش تقييم ساكن قوي للفعل مع تنفيذ تفاعلي أضعف كثيرًا. وكثيرًا ما تبدأ الإخفاقات قبل التنفيذ: إذ يتوقف الوكلاء بعد استقصاء غير مكتمل أو يتصرفون قبل ثبوت الدليل المطلوب. وبمجرد الحصول على الدليل المطلوب، يكون تنفيذ الإجراء الواحد موثوقًا عادةً، بينما تكشف سير العمل متعددة الإجراءات أيضًا عن متطلبات مسبقة غير مستوفاة وتنفيذ غير مكتمل. لهذا التحليل، نقدم SAFE-ACTBENCH، الذي يضم 656 حالة عبر ستة مجالات تشغيلية وخمسة بروتوكولات تتدرج من الحكم الساكن على الفعل وعدم الفعل بعد الاستقصاء إلى سير عمل أحادية ومتعددة الإجراءات. ويتتبع سجل أدلة مرتبط بالمصدر ومُقيِّم مسارات قطعي المعلومات التي أُثبتت ووقت حدوث الإجراءات وما إذا كانت التبعيات اللاحقة مستوفاة. وتُظهر هذه النتائج أن الإخفاقات لا تنشأ عن نقص المعلومات فحسب، بل أيضًا عن كيفية استخدام الوكلاء للأدلة المثبتة عند اتخاذ القرار وتنفيذ الإجراءات.
One-sentence Summary
Researchers from the National University of Singapore, Hong Kong Baptist University, Amazon Web Services, and Princeton University introduce SAFE-ACTBENCH, a 656-case benchmark across six operational domains and five protocols with a provenance-bound Evidence Ledger and deterministic trajectory evaluator, showing that tool-using agent failures arise not only from missing information but also from how agents use established evidence when deciding and executing actions.
Key Contributions
- The paper formulates consequential tool use as an observable evidence-to-action chain linking investigation, execution, and downstream dependencies in interactive environments.
- It introduces SAFE-ACTBENCH, a benchmark suite of 656 cases across six operational domains and five protocols, along with a provenance-bound Evidence Ledger and deterministic trajectory evaluator that verify required evidence, tool calls, targets, arguments, action dependencies, and resulting state changes.
- Results across ten model-harness configurations characterize where this chain breaks: strong static action assessment can coexist with weaker interactive execution, failures often arise from incomplete investigation or premature action, and multi-action workflows expose unresolved prerequisites and incomplete execution.
Introduction
LLM agents increasingly perform consequential tool calls, such as issuing refunds or updating records, where correctness depends not only on choosing the right action but also on establishing task-relevant evidence before execution. Prior evaluation has moved toward interactive, multi-step environments, but it rarely checks whether each state-changing action is observably supported by evidence bound to the correct entity and prior results. The authors formulate this problem as an evidence-to-action chain linking investigation, execution, and downstream dependencies. They introduce SAFEACTBENCH, a benchmark suite with 656 cases across six operational domains and deterministic trajectory evaluation that localizes failures in evidence gathering, action binding, timing, and multi-step execution.
Dataset
The authors introduce SAFEACTBENCH as a benchmark for evidence-to-action reasoning in tool-using tasks. It contains 656 cases across six operational domains. Each case specifies the evidence required before an action, the action that becomes supported, and the expected result or state change.
-
Sources and composition:
- Customer and policy operations: 112 cases
- Engineering and infrastructure operations: 109 cases
- Legal and financial operations: 144 cases
- Research assistance: 97 cases
- Smart-home control: 96 cases
- Healthcare operations: 98 cases
-
Protocol structure: The dataset includes one static protocol and four interactive protocols. Legacy requires a structured ALLOW/BLOCK/DEFER judgment for a fixed candidate action. V0 requires investigation to justify non-execution. V1 requires all necessary evidence before exactly one correct consequential action. V2 extends to a linear workflow where later actions consume actual earlier results. V3 allows any topological order consistent with a directed acyclic dependency graph.
-
Case schema: Each case provides task context, information tools, consequential tools, and one or more consequential actions. For each action, the specification records evidence requirements, required operation, target, arguments, and any expected results or state changes. Multi-action dependencies are represented as either a linear sequence or a directed acyclic graph.
-
Processing and verification: Evidence requirements and expected actions are hidden from the agent and used only for evaluation. A prerequisite is established only when the task context or prior observations provide valid evidence with correct provenance. A deterministic evaluator checks each case against its specification. The primary score is binary exact case success. The evaluator checks required investigation, correct tools, targets, arguments, produced results, dependency order, and terminal conditions. It does not inspect hidden reasoning or use an LLM judge.
-
Provenance construction: An Evidence Ledger binds established facts to their source interactions and the entities or states they describe. Evidence for one entity does not transfer to another, even if values match. Similarly, an action result may directly support a dependent action only when the required source is the relevant predecessor; otherwise a separate state observation is required.
-
Usage: The data is used as an evaluation benchmark rather than as a training set in the described section. The text does not report a training split or mixture ratios. It states that all 656 author reference solutions pass the deterministic evaluator after human audits of the case specifications and evaluator.
-
Cropping: No cropping strategy is described for this dataset.
Method
The authors propose the evidence-to-action chain framework to evaluate how agents translate environmental information into consequential behavior. This framework distinguishes between information-gathering tool calls and consequential actions that modify task-relevant external state. The core premise is that every consequential action must be supported by evidence established prior to its execution. Endpoint correctness alone is insufficient if the preceding interaction fails to support the intended target action.
To formalize this requirement, the authors introduce a provenance-bound evaluation mechanism. Evidence is strictly bound to the specific entity and state it describes. For instance, querying a different entity that returns matching values does not establish the necessary conditions for the target action.
As shown in the figure below, an agent tasked with refunding a specific charge must first query that exact charge to verify its duplicate status, eligibility, and amount. If the agent queries a different charge and then proceeds to refund the target charge, the action may be permitted by the full task state and succeed at the endpoint level. However, the observed evidence before execution remains unsupported because the information gathered does not establish the entity and state used by the action. The evaluation focuses on observable interaction trajectories rather than internal reasoning, ensuring that the information available before execution strictly justifies the executed action.
The framework also governs investigated non-action and chained actions. When no consequential action should occur, the agent must complete the task-relevant investigation to establish why, and then stop without producing side effects. In multi-action workflows, the results of earlier actions serve as necessary evidence for subsequent steps. Dependencies constrain the timing and execution order, meaning a dependent action can only occur after its required predecessors have completed and the necessary outputs are available. The evidence-to-action chain breaks if evidence is bound to the wrong entity, execution occurs prematurely, or dependencies are violated.
To measure this chain deterministically, the authors design a verification system that checks each case against its specification without inspecting hidden reasoning or relying on an LLM judge. The primary metric, exact case success, is evaluated as binary for each episode. For a consequential action a, the support condition is defined as:
Supported(a)⟺∀r∈R(a),r is established before a.The deterministic evaluator replays trajectories to verify whether the required investigation is completed, whether consequential calls use the correct tools, targets, and arguments, and whether action dependencies and terminal conditions are respected. Failure of any required component renders the case unsuccessful, ensuring that the evaluation strictly relies on task context, observed tool interactions, and the resulting environment state.
Experiment
The evaluation compares ten model-harness configurations across five model families, each paired with its family-associated harness and a shared Inspect ReAct harness on the same public tasks and tool interface. Main results show that strong legacy-task performance can mask weaker evidence-grounded execution, while harness effects are model dependent. Behavioral analyses localize failures upstream at investigation and action timing rather than single-action execution, and controlled interventions show that withholding or contradicting evidence reduces but does not stop action, with agents responding more strongly to requester claims than to absent evidence alone.
SAFEACTBENCH spans one static decision protocol and four interactive execution regimes that progressively require evidence provenance and dependency-aware action sequencing. Most cases belong to interactive protocols, with the investigated non-action setting as the largest category and the other interactive protocols having comparable case counts. Success is scored deterministically against protocol-specific evidence, tool-use, result, and dependency conditions rather than by an LLM judge. Static Legacy evaluation checks only a structured ALLOW/BLOCK/DEFER judgment on a fixed candidate action, while interactive protocols score the full trajectory. Interactive regimes require evidence to be established with correct provenance before each consequential action, and later actions must use actual predecessor results. The investigated non-action protocol is the largest interactive category, while single-action, linear multi-action, and dependency-constrained multi-action protocols are similar in size. A deterministic evaluator assigns binary exact case success using task context, observed tool interactions, and resulting environment state.
Reported models perform strongly on Legacy tasks, but their V1, V2, and V3 exact case success is much lower and more variable, showing that Legacy success does not guarantee evidence-grounded execution. Harness effects are substantial and model-dependent, with some models gaining overall while improving only certain protocols. GLM-ZCode and DeepSeek-DSH both exceed 96% on Legacy, but their V1-V3 results diverge sharply: DeepSeek remains around 60% while GLM falls to 12.1-34.1%. DeepSeek-V4 improves overall from Inspect to DSH, partly through Legacy, while Claude-5 and GPT-5.6 also perform better with their non-Inspect harnesses than with Inspect.
Failures on the benchmark often occur before supported execution. Agents frequently stop before completing required investigation or act before evidence is complete, while action execution after completed evidence is generally reliable. Multi-step workflows show multiple overlapping breakdowns, including unresolved prerequisites and incomplete execution, with their prevalence varying across configurations. Conditional action success after evidence completion is high for most configurations, with DeepSeek-Inspect lower, indicating that execution is not the main bottleneck. Incomplete investigation and premature action are widespread, so failures often occur before a supported action attempt. Unresolved prerequisites and incomplete workflow execution are distinct but overlapping multi-step failure modes, and their relative prevalence differs across configurations.
Paired harness comparisons on the same cases show model-family-specific differences in success: DeepSeek's family-associated harness significantly improves over Inspect, while GLM's family-associated harness significantly lowers success. Claude, GPT, and Qwen show positive point estimates with confidence intervals that include or reach zero. Investigation completion is similar for DeepSeek across harnesses but lower for GLM under its family-associated harness. DeepSeek's family-associated harness gains success over Inspect, while GLM's family-associated harness loses success, with confidence intervals excluding zero in both cases. Claude, GPT, and Qwen have positive point estimates, but their confidence intervals include or reach zero, and GLM's investigation completion is lower under its family-associated harness.
SAFEACTBENCH evaluates agents across one static decision protocol and four interactive execution regimes requiring evidence provenance and dependency-aware action sequencing, with deterministic scoring based on protocol-specific conditions. The experiments show that strong performance on static Legacy tasks does not transfer to evidence-grounded execution, and harness effects are substantial and model-dependent. Most failures occur before a supported action is attempted, mainly from incomplete investigation or premature action rather than from unreliable execution. Paired harness comparisons confirm significant success gains for DeepSeek under its family-associated harness, significant degradation for GLM, and inconclusive gains for Claude, GPT, and Qwen.