Command Palette
Search for a command to run...
証拠から行動へ:ツール利用エージェントはいかに失敗するか
証拠から行動へ:ツール利用エージェントはいかに失敗するか
Hongzhan Lin Shidong Cao Ziyang Luo Wenhao Chai Mong-Li Lee Wynne Hsu
概要
ツール利用エージェントは外部状態に重大な変更を加えるが、結果が正しいことは、その行動が事前に確立された証拠によって裏付けられていたことを保証しない。本研究では、エージェントが行動するか否かの判断から単一行動や依存関係を持つワークフローの実行へと進む過程で、この証拠から行動への連鎖がどこで破綻するかを検討する。10種類のモデル・ハーネス構成では、静的な行動評価が高い性能を示す一方で、対話的な実行ははるかに弱いという共存がみられる。失敗は実行前に始まることが多く、エージェントは調査が不完全なまま停止したり、必要な証拠が確立される前に行動したりする。必要な証拠が得られれば単一行動の実行は通常信頼できるが、複数行動ワークフローでは未解決の前提条件や不完全な実行がさらに顕在化する。この分析のため、我々はSAFE-ACTBENCHを導入する。これは、6つの運用ドメインと、静的な行動判断および調査済み非行動から単一・複数行動ワークフローへと進む5つのプロトコルにわたる656件の事例から構成される。来歴に結びついたEvidence Ledgerと決定的軌跡評価器により、どの情報が確立され、いつ行動が生じ、下流の依存関係が満たされたかを追跡する。これらの結果は、失敗が情報の欠如だけでなく、エージェントが行動の判断と実行の際に確立済みの証拠をどのように用いるかからも生じることを示す。
One-sentence Summary
Researchers from the National University of Singapore, Hong Kong Baptist University, Amazon Web Services, and Princeton University introduce SAFE-ACTBENCH, a 656-case benchmark across six operational domains and five protocols with a provenance-bound Evidence Ledger and deterministic trajectory evaluator, showing that tool-using agent failures arise not only from missing information but also from how agents use established evidence when deciding and executing actions.
Key Contributions
- The paper formulates consequential tool use as an observable evidence-to-action chain linking investigation, execution, and downstream dependencies in interactive environments.
- It introduces SAFE-ACTBENCH, a benchmark suite of 656 cases across six operational domains and five protocols, along with a provenance-bound Evidence Ledger and deterministic trajectory evaluator that verify required evidence, tool calls, targets, arguments, action dependencies, and resulting state changes.
- Results across ten model-harness configurations characterize where this chain breaks: strong static action assessment can coexist with weaker interactive execution, failures often arise from incomplete investigation or premature action, and multi-action workflows expose unresolved prerequisites and incomplete execution.
Introduction
LLM agents increasingly perform consequential tool calls, such as issuing refunds or updating records, where correctness depends not only on choosing the right action but also on establishing task-relevant evidence before execution. Prior evaluation has moved toward interactive, multi-step environments, but it rarely checks whether each state-changing action is observably supported by evidence bound to the correct entity and prior results. The authors formulate this problem as an evidence-to-action chain linking investigation, execution, and downstream dependencies. They introduce SAFEACTBENCH, a benchmark suite with 656 cases across six operational domains and deterministic trajectory evaluation that localizes failures in evidence gathering, action binding, timing, and multi-step execution.
Dataset
The authors introduce SAFEACTBENCH as a benchmark for evidence-to-action reasoning in tool-using tasks. It contains 656 cases across six operational domains. Each case specifies the evidence required before an action, the action that becomes supported, and the expected result or state change.
-
Sources and composition:
- Customer and policy operations: 112 cases
- Engineering and infrastructure operations: 109 cases
- Legal and financial operations: 144 cases
- Research assistance: 97 cases
- Smart-home control: 96 cases
- Healthcare operations: 98 cases
-
Protocol structure: The dataset includes one static protocol and four interactive protocols. Legacy requires a structured ALLOW/BLOCK/DEFER judgment for a fixed candidate action. V0 requires investigation to justify non-execution. V1 requires all necessary evidence before exactly one correct consequential action. V2 extends to a linear workflow where later actions consume actual earlier results. V3 allows any topological order consistent with a directed acyclic dependency graph.
-
Case schema: Each case provides task context, information tools, consequential tools, and one or more consequential actions. For each action, the specification records evidence requirements, required operation, target, arguments, and any expected results or state changes. Multi-action dependencies are represented as either a linear sequence or a directed acyclic graph.
-
Processing and verification: Evidence requirements and expected actions are hidden from the agent and used only for evaluation. A prerequisite is established only when the task context or prior observations provide valid evidence with correct provenance. A deterministic evaluator checks each case against its specification. The primary score is binary exact case success. The evaluator checks required investigation, correct tools, targets, arguments, produced results, dependency order, and terminal conditions. It does not inspect hidden reasoning or use an LLM judge.
-
Provenance construction: An Evidence Ledger binds established facts to their source interactions and the entities or states they describe. Evidence for one entity does not transfer to another, even if values match. Similarly, an action result may directly support a dependent action only when the required source is the relevant predecessor; otherwise a separate state observation is required.
-
Usage: The data is used as an evaluation benchmark rather than as a training set in the described section. The text does not report a training split or mixture ratios. It states that all 656 author reference solutions pass the deterministic evaluator after human audits of the case specifications and evaluator.
-
Cropping: No cropping strategy is described for this dataset.
Method
The authors propose the evidence-to-action chain framework to evaluate how agents translate environmental information into consequential behavior. This framework distinguishes between information-gathering tool calls and consequential actions that modify task-relevant external state. The core premise is that every consequential action must be supported by evidence established prior to its execution. Endpoint correctness alone is insufficient if the preceding interaction fails to support the intended target action.
To formalize this requirement, the authors introduce a provenance-bound evaluation mechanism. Evidence is strictly bound to the specific entity and state it describes. For instance, querying a different entity that returns matching values does not establish the necessary conditions for the target action.
As shown in the figure below, an agent tasked with refunding a specific charge must first query that exact charge to verify its duplicate status, eligibility, and amount. If the agent queries a different charge and then proceeds to refund the target charge, the action may be permitted by the full task state and succeed at the endpoint level. However, the observed evidence before execution remains unsupported because the information gathered does not establish the entity and state used by the action. The evaluation focuses on observable interaction trajectories rather than internal reasoning, ensuring that the information available before execution strictly justifies the executed action.
The framework also governs investigated non-action and chained actions. When no consequential action should occur, the agent must complete the task-relevant investigation to establish why, and then stop without producing side effects. In multi-action workflows, the results of earlier actions serve as necessary evidence for subsequent steps. Dependencies constrain the timing and execution order, meaning a dependent action can only occur after its required predecessors have completed and the necessary outputs are available. The evidence-to-action chain breaks if evidence is bound to the wrong entity, execution occurs prematurely, or dependencies are violated.
To measure this chain deterministically, the authors design a verification system that checks each case against its specification without inspecting hidden reasoning or relying on an LLM judge. The primary metric, exact case success, is evaluated as binary for each episode. For a consequential action a, the support condition is defined as:
Supported(a)⟺∀r∈R(a),r is established before a.The deterministic evaluator replays trajectories to verify whether the required investigation is completed, whether consequential calls use the correct tools, targets, and arguments, and whether action dependencies and terminal conditions are respected. Failure of any required component renders the case unsuccessful, ensuring that the evaluation strictly relies on task context, observed tool interactions, and the resulting environment state.
Experiment
The evaluation compares ten model-harness configurations across five model families, each paired with its family-associated harness and a shared Inspect ReAct harness on the same public tasks and tool interface. Main results show that strong legacy-task performance can mask weaker evidence-grounded execution, while harness effects are model dependent. Behavioral analyses localize failures upstream at investigation and action timing rather than single-action execution, and controlled interventions show that withholding or contradicting evidence reduces but does not stop action, with agents responding more strongly to requester claims than to absent evidence alone.
SAFEACTBENCH spans one static decision protocol and four interactive execution regimes that progressively require evidence provenance and dependency-aware action sequencing. Most cases belong to interactive protocols, with the investigated non-action setting as the largest category and the other interactive protocols having comparable case counts. Success is scored deterministically against protocol-specific evidence, tool-use, result, and dependency conditions rather than by an LLM judge. Static Legacy evaluation checks only a structured ALLOW/BLOCK/DEFER judgment on a fixed candidate action, while interactive protocols score the full trajectory. Interactive regimes require evidence to be established with correct provenance before each consequential action, and later actions must use actual predecessor results. The investigated non-action protocol is the largest interactive category, while single-action, linear multi-action, and dependency-constrained multi-action protocols are similar in size. A deterministic evaluator assigns binary exact case success using task context, observed tool interactions, and resulting environment state.
Reported models perform strongly on Legacy tasks, but their V1, V2, and V3 exact case success is much lower and more variable, showing that Legacy success does not guarantee evidence-grounded execution. Harness effects are substantial and model-dependent, with some models gaining overall while improving only certain protocols. GLM-ZCode and DeepSeek-DSH both exceed 96% on Legacy, but their V1-V3 results diverge sharply: DeepSeek remains around 60% while GLM falls to 12.1-34.1%. DeepSeek-V4 improves overall from Inspect to DSH, partly through Legacy, while Claude-5 and GPT-5.6 also perform better with their non-Inspect harnesses than with Inspect.
Failures on the benchmark often occur before supported execution. Agents frequently stop before completing required investigation or act before evidence is complete, while action execution after completed evidence is generally reliable. Multi-step workflows show multiple overlapping breakdowns, including unresolved prerequisites and incomplete execution, with their prevalence varying across configurations. Conditional action success after evidence completion is high for most configurations, with DeepSeek-Inspect lower, indicating that execution is not the main bottleneck. Incomplete investigation and premature action are widespread, so failures often occur before a supported action attempt. Unresolved prerequisites and incomplete workflow execution are distinct but overlapping multi-step failure modes, and their relative prevalence differs across configurations.
Paired harness comparisons on the same cases show model-family-specific differences in success: DeepSeek's family-associated harness significantly improves over Inspect, while GLM's family-associated harness significantly lowers success. Claude, GPT, and Qwen show positive point estimates with confidence intervals that include or reach zero. Investigation completion is similar for DeepSeek across harnesses but lower for GLM under its family-associated harness. DeepSeek's family-associated harness gains success over Inspect, while GLM's family-associated harness loses success, with confidence intervals excluding zero in both cases. Claude, GPT, and Qwen have positive point estimates, but their confidence intervals include or reach zero, and GLM's investigation completion is lower under its family-associated harness.
SAFEACTBENCH evaluates agents across one static decision protocol and four interactive execution regimes requiring evidence provenance and dependency-aware action sequencing, with deterministic scoring based on protocol-specific conditions. The experiments show that strong performance on static Legacy tasks does not transfer to evidence-grounded execution, and harness effects are substantial and model-dependent. Most failures occur before a supported action is attempted, mainly from incomplete investigation or premature action rather than from unreliable execution. Paired harness comparisons confirm significant success gains for DeepSeek under its family-associated harness, significant degradation for GLM, and inconclusive gains for Claude, GPT, and Qwen.