HyperAIHyperAI

Command Palette

Search for a command to run...

Agent
LLM

Mid-Harness : mise à l’échelle des actions entre le modèle et le harnais pour les agents terminaux

Résumé

Les agents terminaux agissent par l’intermédiaire de générations stochastiques du modèle, mais la capacité à générer une action utile ne garantit pas son exécution fiable. Une mauvaise commande (par exemple, l’installation d’un mauvais paquet) peut modifier l’environnement de manière à entraver la progression ultérieure, même lorsque le modèle aurait pu générer une meilleure alternative. Nous cherchons à déterminer si l’allocation de calcul au moment du test à la frontière modèle-harnais peut améliorer la fiabilité des actions et le succès des trajectoires, et ce qui rend cette allocation efficace. Pour étudier ces questions, nous présentons Mid-Harness, qui échantillonne et vérifie des actions candidates avant d’en transmettre une pour exécution, tout en laissant inchangés le générateur et le harnais. Avec un générateur TMAX-9B, un échantillonnage accru d’actions n’apporte que peu de bénéfices lorsque la vérification est faible, tandis qu’un vérificateur performant peut exploiter des alternatives utiles issues du même générateur. Sur TerminalBench-Lite, un vérificateur GPT-5.6 Sol fait passer le Pass@1 de 50,00 % pour l’agent de base à 68,03 % avec 8 actions échantillonnées. Lorsque le même modèle TMAX-9B sert de vérificateur, la vérification par paires est la plus performante parmi les mécanismes de vérification évalués. La distillation des réponses du vérificateur plus puissant dans TMAX-9B améliore encore le Pass@1, sans modifier le générateur d’actions. Avec TMAX-9B sur TerminalBench-Lite, la combinaison de la mise à l’échelle des actions et des trajectoires atteint un succès plus élevé pour un coût estimé en jetons inférieur à celui de la seule génération de trajectoires supplémentaires. Mid-Harness améliore également les performances sur d’autres modèles, bancs d’essai et harnais. Ces résultats identifient la mise à l’échelle des actions comme une cible prometteuse pour la mise à l’échelle du calcul au moment du test chez les agents terminaux. La page du projet est disponible à l’adresse link.

One-sentence Summary

Researchers from NVIDIA and KAIST propose Mid-Harness, which samples and verifies candidate actions at the model-harness boundary before execution while keeping the generator and harness unchanged; with a TMAX-9B generator on TerminalBench-Lite, a GPT-5.6 Sol verifier lifts Pass@1 from 50.00% to 68.03% using 8 sampled actions, distilling the stronger verifier's responses into TMAX-9B further improves Pass@1, and combining action and trajectory scaling achieves higher success at lower estimated token cost than generating more trajectories alone.

Key Contributions

  • Mid-Harness samples and verifies candidate actions at the model-harness boundary before forwarding one action for execution, while keeping the generator and harness unchanged.
  • On TerminalBench-Lite, verification quality governs action sampling benefits: a GPT-5.6 Sol verifier with a TMAX-9B generator raises Pass@1 from 50.00% to 68.03% with 8 sampled actions, and pairwise verification performs best among evaluated mechanisms when TMAX-9B serves as verifier.
  • Distilling the stronger verifier into TMAX-9B improves Pass@1 without changing the action generator. Combining Mid-Harness with trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone, with gains across additional models, benchmarks, and harnesses.

Introduction

Large language models increasingly power terminal agents for software engineering, data science, and scientific discovery. A core challenge is action reliability: even a capable model may fail a long-horizon task because each stochastic command changes the environment, and a single poor command can derail later decisions. Prior work improves performance by scaling test-time compute via sampling, verification, and refinement, but it leaves open when and why action-level compute actually improves trajectory success because candidate quality and verification must be studied together. The authors introduce Mid-Harness, which samples several candidate actions per step, verifies them before execution, and executes only the selected action while keeping the generator and harness fixed. This lets them systematically assess candidate width, verification mechanisms, and verifier capability, showing that verification strength largely determines whether action sampling improves task success.

Method

The authors introduce Mid-Harness, a framework designed to study action-level compute scaling at the model-harness boundary while keeping the action generator and execution harness unchanged. At each step of an interaction, the method requests multiple candidate actions from the generator using the same interaction history, applies a verification mechanism, and forwards only the chosen candidate to the harness for execution. This approach allows the system to improve trajectory success without modifying model weights or the serving architecture.

Let π\piπ be the action-generating model and ht=(o0,a1,o1,…,at−1,ot−1)h_t = (o_0, a_1, o_1, \ldots, a_{t-1}, o_{t-1})ht​=(o0​,a1​,o1​,…,at−1​,ot−1​) represent the interaction history, where o0o_0o0​ contains the task instruction. Instead of drawing and executing a single action from π(⋅∣ht)\pi(\cdot \mid h_t)π(⋅∣ht​), Mid-Harness samples NNN candidates conditioned on the same history and identifies the optimal action using a verifier ψ\psiψ:

At=(at1,…,atN),ati∼π(⋅∣ht),at⋆=Verifyψ(ht,At)∈At.\mathcal{A}_t = (a_t^1, \ldots, a_t^N), \quad a_t^i \sim \pi(\cdot \mid h_t), \qquad a_t^\star = \mathrm{Verify}_\psi(h_t, \mathcal{A}_t) \in \mathcal{A}_t.At​=(at1​,…,atN​),ati​∼π(⋅∣ht​),at⋆​=Verifyψ​(ht​,At​)∈At​.

The harness then executes at⋆a_t^\starat⋆​, receives the observation oto_tot​, and updates the history. The remaining candidates are discarded.

To evaluate the proposed actions before execution, the authors compare three distinct verification mechanisms. In listwise verification, the verifier receives the entire candidate set in a single prompt and returns a choice, exposing all alternatives for direct comparison but requiring the model to resolve a complex joint ranking. Pointwise verification decomposes the process into NNN independent evaluations, assigning each candidate a scalar score qi=ψpoint(ht,ati)q_i = \psi_{\mathrm{point}}(h_t, a_t^i)qi​=ψpoint​(ht​,ati​) and selecting the one with the maximum score. This allows for parallel evaluation but struggles to place actions with different purposes on a comparable scale. Pairwise verification focuses on comparing two candidates at a time under the same history, generating comparative scores Jij=ψpair(ht,ati,atj)J_{ij} = \psi_{\mathrm{pair}}(h_t, a_t^i, a_t^j)Jij​=ψpair​(ht​,ati​,atj​). The default pairwise verifier ranks candidates by margin-weighted win rates over the evaluated pairs and returns the highest-ranked action.

The performance of these mechanisms varies significantly with candidate width and verifier capability. As shown in the figure below, pairwise verification generally outperforms listwise and pointwise approaches, and distilling a frontier verifier further enhances performance.

To bridge the quality gap between a frontier verifier and the generator model acting as a verifier, the authors employ a verifier distillation process. They collect pairwise responses from a strong frontier model across numerous trajectories. Using this dataset, they apply Low-Rank Adaptation to fine-tune the verifier to generate the teacher model's reasoning, scores, and preference labels. This fine-tuning is exclusively activated for the verifier, leaving the generator entirely unchanged. This distillation process effectively transfers verification capabilities, raising trajectory success rates without requiring the frontier model to be served at inference time.

Experiment

The evaluation uses TerminalBench Lite with TMAX models under the Vanillux2 harness, comparing action-level sampling and verification with trajectory-level methods through Pass@1 and Pass@3. The experiments show that sampled candidates contain useful alternatives that self-verification only partially recovers, with pairwise verification performing best and distillation from a frontier verifier improving agreement and trajectory success. Action scaling composes with Best-of-N and Sequential Refine, transfers across model scales, benchmarks, harnesses, and non-TMAX models, while remaining verification errors center on command semantics and execution feasibility.

Distillation improves alignment between the verifier and the frontier verifier on the offline benchmark. Score mean absolute error drops by more than half, while pairwise agreement and verification agreement both rise substantially. The distilled verifier is therefore more likely to select the same candidate as the frontier verifier. Distilled verification reduces score error by more than half compared with zero-shot verification. Pairwise agreement rises by about 15 percentage points and verification agreement by about 19 percentage points after distillation.

Composing action verification with trajectory scaling improves terminal task success across model sizes. For TMAX-9B, adding zero-shot or distilled Mid-Harness to Best-of trajectory generation raises Pass@1 from 55.10% to 61.22% or 66.33%, while combining distilled Mid-Harness with sequential refinement raises Pass@1 to 60.20% and Pass@3 to 75.51%. Action verification benefits generators from 4B to 27B, with distilled verification generally giving the largest gains. Zero-shot action verification alone improves Pass@1 over the base agent at all tested model scales, with gains of roughly two to five points. Best-of trajectory selection alone is mixed: it improves Pass@1 for 9B and 27B but slightly reduces it for 4B. Composing zero-shot action verification with Best-of trajectory generation yields strong Pass@1 gains, especially for 9B and 27B. Distilled action verification further improves composition, reaching 66.33% Pass@1 with Best-of at 9B and 60.20% Pass@1 with sequential refinement.

Mid-Harness action verification transfers across multiple benchmarks, models, and harnesses, with zero-shot verification improving Pass@3 in every reported setting and matching or improving Pass@1. Gains are especially pronounced on FeatureBench-Mini, where the base agent rarely succeeds but zero-shot and distilled verification raise success substantially. The distilled variant improves TMAX-9B on SWE-bench-Verified and FeatureBench-Mini Pass@1, though it can trail zero-shot verification on FeatureBench-Mini Pass@3 and does not improve over it on Terminal-Bench 2.1. Zero-shot Mid-Harness improves Pass@3 across all reported settings and matches or improves Pass@1, including for models outside the TMAX family and with a different harness. On FeatureBench-Mini, TMAX-9B starts with very low success, and both zero-shot and distilled verification produce large Pass@1 and Pass@3 gains, with distilled verification raising Pass@1 further but achieving lower Pass@3 than zero-shot. Distilled Mid-Harness improves TMAX-9B on SWE-bench-Verified for both Pass@1 and Pass@3, while on Terminal-Bench 2.1 it matches Pass@3 but does not improve Pass@1 over zero-shot.

The experiments evaluate Mid-Harness action verification in three settings. Distillation substantially aligns the verifier with a frontier verifier offline, reducing score error and improving agreement metrics. Composing action verification with trajectory scaling improves terminal task success across model sizes, with distilled verification usually providing the largest gains. The method also transfers across benchmarks, models, and harnesses, where zero-shot verification improves Pass@3 in every reported setting and is especially helpful on tasks with very low base success, while distilled verification improves some additional settings but occasionally trails zero-shot.


Créer de l'IA avec l'IA

De l'idée au lancement — accélérez votre développement IA avec le co-codage IA gratuit, un environnement prêt à l'emploi et le meilleur prix pour les GPU.

Codage assisté par IA
GPU prêts à l’emploi
Tarifs les plus avantageux

HyperAI Newsletters

Abonnez-vous à nos dernières mises à jour
Nous vous enverrons les dernières mises à jour de la semaine dans votre boîte de réception à neuf heures chaque lundi matin
Propulsé par MailChimp