HyperAIHyperAI

Command Palette

Search for a command to run...

Mise à l'échelle des trajectoires pour des tâches complexes par réécriture auto-récursive

Zongxia Li Yucheng Shi Zhongzhi Li Junyao Yang Ruhan Wang Chengsong Huang Fuxiao Liu Haitao Mi Jordan Boyd-Graber Leowei Liang

Résumé

Les trajectoires réussies sur des tâches difficiles constituent une supervision précieuse pour l'amélioration des modèles. Différents harnais permettent au même modèle de résoudre ces tâches de différentes manières. Leurs trajectoires réussies fournissent une expérience utile pour l'auto-amélioration, mais elles contiennent aussi des interventions du contrôleur et des conventions de flux de travail qui peuvent être indisponibles sous un harnais général. Nous proposons Recursive Self-Rewrite (RSR), un cadre qui convertit ces expériences en capacités de modèle réutilisables par réécriture auto-récursive des trajectoires. Nous utilisons un modèle de base unique, Qwen-3.8-27B, pour découvrir des solutions réussies sous divers harnais et les réécrire en vue d'un apprentissage sous un harnais général. RSR contient un planificateur qui extrait les procédures utiles dans des runbooks, un critique qui détecte les fuites du vérificateur et de la solution, rejette les candidats et les régénère de manière récursive à l'aide du retour critique, et un exécuteur qui suit les runbooks qualifiés pour résoudre chaque tâche dans un nouveau bac à sable sous un harnais général. Nous montrons que l'utilisation de trois harnais élargit la gamme de domaines de tâches que Qwen-3.8-27B peut résoudre et produit des trajectoires réussies plus précieuses ; avec RSR, nous reconstruisons en outre celles-ci en un plus grand ensemble de trajectoires de haute qualité. Nous collectons de l'expérience à partir d'environ 3 000 tâches de terminal auto-constituées. Sur 3 000 tâches, l'union des trois harnais résout 759 tâches, soit 34,3 % de plus que le harnais individuel le plus performant du pool enregistré. Nous utilisons en outre RSR pour réécrire les trajectoires sources réussies, faisant passer l'ensemble d'entraînement de 2 001 à 11 094 trajectoires de haute qualité pour l'ajustement fin de Qwen-3.8-27B. L'entraîner sur ces trajectoires réécrites surpasse à la fois le modèle de base et la SFT directe sur trajectoires : le pass@3 augmente de 57,0 % à 74,2 % sur Terminal-Bench 2, de 1,5 % à 9,1 % sur Terminal-Bench 4, de 39,0 % à 63,0 % sur notre Terminal-Bench Hard auto-constitué, et de 3,0 % à 6,0 % sur notre Software Terminal-Bench. Sur Long-Horizon Terminal-Bench, la récompense de processus passe de 0,21 à 0,29.

One-sentence Summary

Recursive Self-Rewrite (RSR), proposed by researchers from Tencent HY LLM Frontier, the University of Maryland, College Park, and other institutions, recursively rewrites successful trajectories from diverse harnesses into reusable runbooks via a planner, a leakage-screening critic, and an executor, expanding Qwen-3.8-27B finetuning data from 2,001 to 11,094 trajectories and improving Terminal-Bench 2 pass@3 from 57.0% to 74.2%.

Key Contributions

  • The paper introduces Recursive Self-Rewrite, a framework that converts successful trajectories from diverse specialized harnesses into verified demonstrations under a general harness. It consists of a planner that extracts reusable runbooks, a critic that screens for verifier and solution leakage and recursively regenerates rejected candidates, and an executor that follows qualified runbooks in fresh sandboxes.
  • It shows that using multiple harnesses as discovery tools broadens task coverage: across approximately 3,000 self-curated terminal tasks, the union of three harnesses solves 759 tasks, 34.3% more than the strongest individual harness.
  • It demonstrates that rewriting these trajectories expands the training set from 2,001 to 11,094 high-quality examples and improves fine-tuned Qwen-3.8-27B over the base model and direct trajectory SFT. Pass@3 rises from 57.0% to 74.2% on Terminal-Bench 2, from 1.5% to 9.1% on Terminal-Bench 4, from 39.0% to 63.0% on Terminal-Bench Hard, and from 3.0% to 6.0% on Software Terminal-Bench, while process reward rises from 0.21 to 0.29 on Long-Horizon Terminal-Bench.

Introduction

The authors study self-improvement for AI agents on difficult terminal tasks, where performance depends not only on the model but also on its execution harness, the system that controls observations, tool use, verification, recovery, and stopping. Prior work shows that specialized harnesses can improve task success without changing model weights, and that different harnesses solve complementary subsets of tasks. However, trajectories collected from these harnesses mix harness-specific controller interventions, prompts, workflow logic, and stopping rules with the underlying problem-solving behavior. Naively training on such data can cause the model to depend on external control patterns that are unavailable under a general harness at inference time. To address this, the authors propose Recursive Self-Rewrite, which collects successful multi-harness trajectories and rewrites them into verified trajectories under a general harness using the same base model in planner, critic, and executor roles. This enables supervised finetuning from the model's own harness-assisted experience and improves Qwen-3.8-27B across several terminal-task benchmarks.

Dataset

The authors describe the terminal-task dataset as follows:

  • Composition and sources: The training task pool is built from two sources. SWR is a self-constructed collection of about 2,500 terminal tasks across 50 domains. The authors also include 420 filtered and modified tasks from the RST dataset.
  • Domain coverage: SWR spans software usage, biology, chemistry, physics, hardware, operations, and security.
  • Scale: Together, these sources define a pool of roughly 3,000 tasks.
  • Filtering and modification: The RST subset is filtered and modified to have increased difficulty.
  • Usage: The pool is used for terminal-task training, where the model must select tools, reason about environment feedback, and carry out task-specific procedures.
  • Additional processing details: The provided text does not specify schema, cropping, metadata construction, or mixture ratios.

Method

The authors propose Recursive Self-Rewrite, a method designed to improve a model under a general-purpose harness by learning from successful trajectories discovered under diverse specialized harnesses. The core idea is that different harnesses help the same base model solve different difficult tasks, and those successful solutions can be rewritten into demonstrations compatible with a single general harness. The overall process is illustrated in the framework diagram below.

The method operates through three primary stages: Multi-Harness Discovery, Trajectory Rewriting, and Verification.

Multi-Harness Discovery The authors utilize multiple discovery harnesses that differ in how they provide control and support during execution, such as progress tracking, continuation, validation, state management, or recovery from failure. Because different harnesses can be effective on different types of tasks, they expand the range of successful solutions beyond what the model could achieve with just one harness. The base model is run under these multiple discovery harnesses to collect successful trajectories.

Trajectory Rewriting for Experience Learning Successful trajectories collected under different harnesses record how the model solves tasks with different workflow logic. The goal is to transform these successful solutions into learning experiences that the model can practice and learn from under a general harness. Trajectory rewriting involves three distinct model roles: a planner, a critic, and an executor.

  • Planner: The planner reconstructs a runbook for each successful source trajectory, providing a structured description of how the task was solved. Before planning, the source trajectory is compacted by retaining the task instruction, the model’s actions, and the environment’s observations, while removing harness-specific control messages. The runbook summarizes the required end state, key milestones, useful checks, recovery strategies, and common pitfalls. To reduce direct answer transfer, runbooks describe the task, relevant interfaces, and validation procedures without directly providing the finished deliverable. Multiple runbook candidates are sampled for each source trajectory.
  • Critic: The model itself acts as a critic to filter candidate runbooks before they are used for execution. Deterministic checks are first applied, such as schema validation, removal of known artifacts, and rejection of unsupported tool references. Then, a model-based critic that sees only the public task instruction and the candidate runbook determines whether the runbook provides useful procedure or leaks information that the executor should not receive. Only runbooks that pass this screening stage are retained for rewriting.
  • Executor: For each approved runbook, the executor re-solves the task in a fresh sandbox under the general harness. The runbook is provided as private guidance during generation but is never written into the public trajectory. The executor must therefore produce a new trajectory based on the current environment rather than replaying the source trajectory.

Trajectory Filtering and Verification The model itself is used to flag values that appear in a trajectory but cannot be derived from the task or the environment, which indicates hidden answer transfer. Demonstrations containing such values are discarded. For finetuning, only the public interaction history, comprising the task instruction, environment observations, and the model’s responses, is kept, while the runbook and critic conversation are removed. At inference time, the finetuned model runs under the general-purpose harness alone, without the source harnesses or private runbooks, ensuring that any planning, checking, recovery, or continuation comes from the model itself.

Case Study Examples The authors examine rewritten tasks to demonstrate how a source experience can be transformed into new trajectories under the general harness. As shown in the figure below, the examples illustrate a passing and a failing execution guided by the same runbook in each case, alongside representative command excerpts from the source and both rewrites.

Experiment

The experiments evaluate a self-improvement pipeline that collects successful terminal-task rollouts from three complementary execution harnesses, Terminus 2, StateM, and Recursive Self-Reflect Terminus, then reconstructs those experiences into standardized trajectories under the general harness. They show that different harnesses unlock different problem-solving behaviors and jointly expand task coverage, while direct fine-tuning on source rollouts produces mixed gains and can introduce looping failures. Rewriting successful trajectories through recursive self-rewrite leads to cleaner, more generalizable training data and improves performance over both the base model and direct fine-tuning, with additional partial progress on long-horizon tasks.

Combining rollouts from Terminus 2, Recursive Self-Reflect Terminus, and StateM expands solved-task coverage beyond any single harness or pair. The three-harness union solves the most tasks overall and adds a substantial number over the strongest individual harness, with relative gains on both RST and SWR. Individual harnesses show complementary strengths, as leaders vary across benchmark splits. The three-harness union yields the highest overall task coverage and a roughly one-third relative increase over the strongest individual harness. Across individual harnesses, Recursive Self-Reflect Terminus leads on RST, while Terminus 2 leads on SWR and the pooled set; pairwise unions consistently exceed any single harness.

Combining rollouts from Terminus 2, Recursive Self-Reflect Terminus, and StateM broadens discovery coverage beyond any single harness. The pooled collection contains 2,001 successful trajectories covering 759 distinct tasks, which is more solved tasks than the strongest individual harness. Per-rollout success rates are similar across harnesses, indicating that multi-harness pooling mainly increases task coverage. The strongest individual harness solves 565 tasks, while the pooled union solves 759 tasks, adding 194 solved tasks and a relative coverage increase of 34.3%. Recursive Self-Reflect Terminus achieves the highest individual rollout success rate at 15.9%, StateM has the lowest at 11.7%, and the pooled set reaches 13.7%.

Across passing trajectories, the three harnesses show distinct execution styles. Terminus 2 produces shorter trajectories with more exploration-oriented commands and more passing rollouts, while StateM produces longer trajectories with more commands per turn but less exploration. RSRT sits between them on trajectory length and exploration, with the most completion claims per trajectory and a small share of passes occurring after rejection. Terminus 2 has the largest number of passing trajectories and the shortest average and median turn counts among the three harnesses. StateM shows the longest median trajectory length and highest commands per turn, but the lowest exploration command share. RSRT records the most completion claims per trajectory and is the only harness with a notable percentage of passes after rejection.

Rewrites guided by the same runbook tend to have more similar command-level behavior than rewrites using different runbooks. This pattern is stronger for Markdown, while OpenFOAM shows a weaker effect for command metrics. Workflow-level similarity is less consistent and varies by task. Same-runbook pairs show higher exact command overlap and command-order similarity than different-runbook pairs, with a clearer gap for Markdown. Tool-transition and action-sequence similarity are not consistently higher within runbooks; OpenFOAM shows little tool-transition difference and slightly lower within-runbook action-sequence similarity.

RSR outperforms both the base model and Direct SFT across the reported benchmarks, with higher pass@3, higher mean per-run pass rate, and better process reward on LHTB. Direct SFT shows mixed results, improving on several benchmarks but declining on TB2, where training without rewriting can introduce looping behaviors. The results suggest that scaling and standardizing successful experiences under a general harness strengthens generalization. RSR achieves the highest pass@3 and mean per-run pass rate across all five reported benchmarks. Direct SFT has mixed results, improving over the base model on TBH, TB3, and TB4 but declining on TB2 and falling below its base mean per-run pass rate.

The experiments evaluate multi-harness rollouts from Terminus 2, Recursive Self-Reflect Terminus, and StateM, along with runbook-guided rewrite similarity and a comparison of RSR against base and Direct SFT models. Combining the three harnesses yields complementary task coverage and the largest set of solved tasks, while the harnesses show distinct execution styles and same-runbook rewrites are more similar at the command level, especially for Markdown. RSR consistently outperforms the base model and Direct SFT, whereas Direct SFT improves on several benchmarks but can degrade on others through looping behaviors.


Créer de l'IA avec l'IA

De l'idée au lancement — accélérez votre développement IA avec le co-codage IA gratuit, un environnement prêt à l'emploi et le meilleur prix pour les GPU.

Codage assisté par IA
GPU prêts à l’emploi
Tarifs les plus avantageux

HyperAI Newsletters

Abonnez-vous à nos dernières mises à jour
Nous vous enverrons les dernières mises à jour de la semaine dans votre boîte de réception à neuf heures chaque lundi matin
Propulsé par MailChimp