Command Palette
Search for a command to run...
Mid-Harness: ターミナルエージェントのためのモデルとハーネス間におけるアクションのスケーリング
Mid-Harness: ターミナルエージェントのためのモデルとハーネス間におけるアクションのスケーリング
概要
ターミナルエージェントは確率的なモデル生成を通じて行動するが、有用なアクションを生成できるとしても、そのアクションが確実に実行されるとは限らない。不適切なコマンド(例えば誤ったパッケージのインストール)は、モデルがより良い代替案を生成できた場合でも、その後の進捗を妨げる形で環境を変化させうる。我々は、モデルとハーネスの境界にテスト時計算を割り当てることがアクションの信頼性と軌跡の成功率を向上させるか、また何がその割り当てを効果的にするのかを調査する。これらの問いを検討するため、生成器とハーネスは変更せずに、実行に移す前に候補アクションをサンプリングし検証するMid-Harnessを提案する。TMAX-9B生成器では、弱い検証の下ではアクションのサンプリングを増やしてもほとんど利益がない一方、有能な検証器は同じ生成器から有用な代替案を活用できる。TerminalBench-Liteにおいて、GPT-5.6 Sol検証器は、8個のサンプルアクションを用いることでベースエージェントのPass@1を50.00%から68.03%へ向上させる。同じTMAX-9Bモデルを検証器として用いる場合、評価した検証機構の中ではペアワイズ検証が最も高い性能を示す。より強力な検証器の応答をTMAX-9Bへ蒸留すると、アクション生成器を変更せずにPass@1がさらに向上する。TMAX-9BをTerminalBench-Liteで用いた場合、アクションのスケーリングと軌跡のスケーリングを組み合わせることで、軌跡を単に増やすよりも低い推定トークンコストで高い成功率に達する。Mid-Harnessは、追加のモデル、ベンチマーク、ハーネスでも性能を向上させる。これらの結果は、ターミナルエージェントにおけるテスト時計算スケーリングの有望な対象としてアクションスケーリングを位置づける。プロジェクトページはlinkで入手可能である。
One-sentence Summary
Researchers from NVIDIA and KAIST propose Mid-Harness, which samples and verifies candidate actions at the model-harness boundary before execution while keeping the generator and harness unchanged; with a TMAX-9B generator on TerminalBench-Lite, a GPT-5.6 Sol verifier lifts Pass@1 from 50.00% to 68.03% using 8 sampled actions, distilling the stronger verifier's responses into TMAX-9B further improves Pass@1, and combining action and trajectory scaling achieves higher success at lower estimated token cost than generating more trajectories alone.
Key Contributions
- Mid-Harness samples and verifies candidate actions at the model-harness boundary before forwarding one action for execution, while keeping the generator and harness unchanged.
- On TerminalBench-Lite, verification quality governs action sampling benefits: a GPT-5.6 Sol verifier with a TMAX-9B generator raises Pass@1 from 50.00% to 68.03% with 8 sampled actions, and pairwise verification performs best among evaluated mechanisms when TMAX-9B serves as verifier.
- Distilling the stronger verifier into TMAX-9B improves Pass@1 without changing the action generator. Combining Mid-Harness with trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone, with gains across additional models, benchmarks, and harnesses.
Introduction
Large language models increasingly power terminal agents for software engineering, data science, and scientific discovery. A core challenge is action reliability: even a capable model may fail a long-horizon task because each stochastic command changes the environment, and a single poor command can derail later decisions. Prior work improves performance by scaling test-time compute via sampling, verification, and refinement, but it leaves open when and why action-level compute actually improves trajectory success because candidate quality and verification must be studied together. The authors introduce Mid-Harness, which samples several candidate actions per step, verifies them before execution, and executes only the selected action while keeping the generator and harness fixed. This lets them systematically assess candidate width, verification mechanisms, and verifier capability, showing that verification strength largely determines whether action sampling improves task success.
Method
The authors introduce Mid-Harness, a framework designed to study action-level compute scaling at the model-harness boundary while keeping the action generator and execution harness unchanged. At each step of an interaction, the method requests multiple candidate actions from the generator using the same interaction history, applies a verification mechanism, and forwards only the chosen candidate to the harness for execution. This approach allows the system to improve trajectory success without modifying model weights or the serving architecture.
Let π be the action-generating model and ht=(o0,a1,o1,…,at−1,ot−1) represent the interaction history, where o0 contains the task instruction. Instead of drawing and executing a single action from π(⋅∣ht), Mid-Harness samples N candidates conditioned on the same history and identifies the optimal action using a verifier ψ:
At=(at1,…,atN),ati∼π(⋅∣ht),at⋆=Verifyψ(ht,At)∈At.The harness then executes at⋆, receives the observation ot, and updates the history. The remaining candidates are discarded.
To evaluate the proposed actions before execution, the authors compare three distinct verification mechanisms. In listwise verification, the verifier receives the entire candidate set in a single prompt and returns a choice, exposing all alternatives for direct comparison but requiring the model to resolve a complex joint ranking. Pointwise verification decomposes the process into N independent evaluations, assigning each candidate a scalar score qi=ψpoint(ht,ati) and selecting the one with the maximum score. This allows for parallel evaluation but struggles to place actions with different purposes on a comparable scale. Pairwise verification focuses on comparing two candidates at a time under the same history, generating comparative scores Jij=ψpair(ht,ati,atj). The default pairwise verifier ranks candidates by margin-weighted win rates over the evaluated pairs and returns the highest-ranked action.
The performance of these mechanisms varies significantly with candidate width and verifier capability. As shown in the figure below, pairwise verification generally outperforms listwise and pointwise approaches, and distilling a frontier verifier further enhances performance.
To bridge the quality gap between a frontier verifier and the generator model acting as a verifier, the authors employ a verifier distillation process. They collect pairwise responses from a strong frontier model across numerous trajectories. Using this dataset, they apply Low-Rank Adaptation to fine-tune the verifier to generate the teacher model's reasoning, scores, and preference labels. This fine-tuning is exclusively activated for the verifier, leaving the generator entirely unchanged. This distillation process effectively transfers verification capabilities, raising trajectory success rates without requiring the frontier model to be served at inference time.
Experiment
The evaluation uses TerminalBench Lite with TMAX models under the Vanillux2 harness, comparing action-level sampling and verification with trajectory-level methods through Pass@1 and Pass@3. The experiments show that sampled candidates contain useful alternatives that self-verification only partially recovers, with pairwise verification performing best and distillation from a frontier verifier improving agreement and trajectory success. Action scaling composes with Best-of-N and Sequential Refine, transfers across model scales, benchmarks, harnesses, and non-TMAX models, while remaining verification errors center on command semantics and execution feasibility.
Distillation improves alignment between the verifier and the frontier verifier on the offline benchmark. Score mean absolute error drops by more than half, while pairwise agreement and verification agreement both rise substantially. The distilled verifier is therefore more likely to select the same candidate as the frontier verifier. Distilled verification reduces score error by more than half compared with zero-shot verification. Pairwise agreement rises by about 15 percentage points and verification agreement by about 19 percentage points after distillation.
Composing action verification with trajectory scaling improves terminal task success across model sizes. For TMAX-9B, adding zero-shot or distilled Mid-Harness to Best-of trajectory generation raises Pass@1 from 55.10% to 61.22% or 66.33%, while combining distilled Mid-Harness with sequential refinement raises Pass@1 to 60.20% and Pass@3 to 75.51%. Action verification benefits generators from 4B to 27B, with distilled verification generally giving the largest gains. Zero-shot action verification alone improves Pass@1 over the base agent at all tested model scales, with gains of roughly two to five points. Best-of trajectory selection alone is mixed: it improves Pass@1 for 9B and 27B but slightly reduces it for 4B. Composing zero-shot action verification with Best-of trajectory generation yields strong Pass@1 gains, especially for 9B and 27B. Distilled action verification further improves composition, reaching 66.33% Pass@1 with Best-of at 9B and 60.20% Pass@1 with sequential refinement.
Mid-Harness action verification transfers across multiple benchmarks, models, and harnesses, with zero-shot verification improving Pass@3 in every reported setting and matching or improving Pass@1. Gains are especially pronounced on FeatureBench-Mini, where the base agent rarely succeeds but zero-shot and distilled verification raise success substantially. The distilled variant improves TMAX-9B on SWE-bench-Verified and FeatureBench-Mini Pass@1, though it can trail zero-shot verification on FeatureBench-Mini Pass@3 and does not improve over it on Terminal-Bench 2.1. Zero-shot Mid-Harness improves Pass@3 across all reported settings and matches or improves Pass@1, including for models outside the TMAX family and with a different harness. On FeatureBench-Mini, TMAX-9B starts with very low success, and both zero-shot and distilled verification produce large Pass@1 and Pass@3 gains, with distilled verification raising Pass@1 further but achieving lower Pass@3 than zero-shot. Distilled Mid-Harness improves TMAX-9B on SWE-bench-Verified for both Pass@1 and Pass@3, while on Terminal-Bench 2.1 it matches Pass@3 but does not improve Pass@1 over zero-shot.
The experiments evaluate Mid-Harness action verification in three settings. Distillation substantially aligns the verifier with a frontier verifier offline, reducing score error and improving agreement metrics. Composing action verification with trajectory scaling improves terminal task success across model sizes, with distilled verification usually providing the largest gains. The method also transfers across benchmarks, models, and harnesses, where zero-shot verification improves Pass@3 in every reported setting and is especially helpful on tasks with very low base success, while distilled verification improves some additional settings but occasionally trails zero-shot.