HyperAIHyperAI

Command Palette

Search for a command to run...

再帰的自己書き換え(Recursive Self-Rewrite)による複雑タスク向け軌跡のスケーリング

Zongxia Li Yucheng Shi Zhongzhi Li Junyao Yang Ruhan Wang Chengsong Huang Fuxiao Liu Haitao Mi Jordan Boyd-Graber Leowei Liang

概要

困難なタスクにおける成功軌跡は、モデル改善のための価値ある教師信号である。異なるハーネスは同じモデルがそれらのタスクを異なる方法で解くことを可能にする。その成功軌跡は自己改善に有用な経験を提供するが、一般的なハーネスでは利用できないコントローラ介入やワークフロー上の慣習も含んでいる。我々は、再帰的な軌跡自己書き換えを通じてこれらの経験を再利用可能なモデル能力へ変換するフレームワークである Recursive Self-Rewrite (RSR) を提案する。単一のベースモデル Qwen-3.8-27B を用い、多様なハーネスの下で成功解を発見し、一般的なハーネスの下での学習のために書き換える。RSR は、有用な手順を runbook(手順書)に抽出する planner、検証器や解答の漏洩を検査し、候補を却下して批評フィードバックを用いて再帰的に再生成する critic、適格な runbook に従って一般的なハーネスの下の新しいサンドボックスで各タスクを解く executor から構成される。3つのハーネスを用いることで、Qwen-3.8-27B が解けるタスク領域が拡大し、より価値ある成功軌跡が得られることを示す。さらに RSR により、これらをより大規模な高品質軌跡セットへ再構成する。約3Kの自己収集した端末タスクから経験を収集する。3Kタスク全体では、3つのハーネスの和集合により759タスクが解け、これは記録されたプール内で最も強い単一ハーネスより34.3%多い。さらに RSR を用いて成功したソース軌跡を書き換え、Qwen-3.8-27B のファインチューニング用に高品質軌跡を2,001件から11,094件へ拡大する。これらの書き換え軌跡で学習すると、ベースモデルおよび直接的な軌跡SFTの両方を上回る。pass@3 は Terminal-Bench 2 で57.0%から74.2%へ、Terminal-Bench 4 で1.5%から9.1%へ、自己収集の Terminal-Bench Hard で39.0%から63.0%へ、Software Terminal-Bench で3.0%から6.0%へ向上する。Long-Horizon Terminal-Bench ではプロセス報酬が0.21から0.29へ上昇する。

One-sentence Summary

Recursive Self-Rewrite (RSR), proposed by researchers from Tencent HY LLM Frontier, the University of Maryland, College Park, and other institutions, recursively rewrites successful trajectories from diverse harnesses into reusable runbooks via a planner, a leakage-screening critic, and an executor, expanding Qwen-3.8-27B finetuning data from 2,001 to 11,094 trajectories and improving Terminal-Bench 2 pass@3 from 57.0% to 74.2%.

Key Contributions

  • The paper introduces Recursive Self-Rewrite, a framework that converts successful trajectories from diverse specialized harnesses into verified demonstrations under a general harness. It consists of a planner that extracts reusable runbooks, a critic that screens for verifier and solution leakage and recursively regenerates rejected candidates, and an executor that follows qualified runbooks in fresh sandboxes.
  • It shows that using multiple harnesses as discovery tools broadens task coverage: across approximately 3,000 self-curated terminal tasks, the union of three harnesses solves 759 tasks, 34.3% more than the strongest individual harness.
  • It demonstrates that rewriting these trajectories expands the training set from 2,001 to 11,094 high-quality examples and improves fine-tuned Qwen-3.8-27B over the base model and direct trajectory SFT. Pass@3 rises from 57.0% to 74.2% on Terminal-Bench 2, from 1.5% to 9.1% on Terminal-Bench 4, from 39.0% to 63.0% on Terminal-Bench Hard, and from 3.0% to 6.0% on Software Terminal-Bench, while process reward rises from 0.21 to 0.29 on Long-Horizon Terminal-Bench.

Introduction

The authors study self-improvement for AI agents on difficult terminal tasks, where performance depends not only on the model but also on its execution harness, the system that controls observations, tool use, verification, recovery, and stopping. Prior work shows that specialized harnesses can improve task success without changing model weights, and that different harnesses solve complementary subsets of tasks. However, trajectories collected from these harnesses mix harness-specific controller interventions, prompts, workflow logic, and stopping rules with the underlying problem-solving behavior. Naively training on such data can cause the model to depend on external control patterns that are unavailable under a general harness at inference time. To address this, the authors propose Recursive Self-Rewrite, which collects successful multi-harness trajectories and rewrites them into verified trajectories under a general harness using the same base model in planner, critic, and executor roles. This enables supervised finetuning from the model's own harness-assisted experience and improves Qwen-3.8-27B across several terminal-task benchmarks.

Dataset

The authors describe the terminal-task dataset as follows:

  • Composition and sources: The training task pool is built from two sources. SWR is a self-constructed collection of about 2,500 terminal tasks across 50 domains. The authors also include 420 filtered and modified tasks from the RST dataset.
  • Domain coverage: SWR spans software usage, biology, chemistry, physics, hardware, operations, and security.
  • Scale: Together, these sources define a pool of roughly 3,000 tasks.
  • Filtering and modification: The RST subset is filtered and modified to have increased difficulty.
  • Usage: The pool is used for terminal-task training, where the model must select tools, reason about environment feedback, and carry out task-specific procedures.
  • Additional processing details: The provided text does not specify schema, cropping, metadata construction, or mixture ratios.

Method

The authors propose Recursive Self-Rewrite, a method designed to improve a model under a general-purpose harness by learning from successful trajectories discovered under diverse specialized harnesses. The core idea is that different harnesses help the same base model solve different difficult tasks, and those successful solutions can be rewritten into demonstrations compatible with a single general harness. The overall process is illustrated in the framework diagram below.

The method operates through three primary stages: Multi-Harness Discovery, Trajectory Rewriting, and Verification.

Multi-Harness Discovery The authors utilize multiple discovery harnesses that differ in how they provide control and support during execution, such as progress tracking, continuation, validation, state management, or recovery from failure. Because different harnesses can be effective on different types of tasks, they expand the range of successful solutions beyond what the model could achieve with just one harness. The base model is run under these multiple discovery harnesses to collect successful trajectories.

Trajectory Rewriting for Experience Learning Successful trajectories collected under different harnesses record how the model solves tasks with different workflow logic. The goal is to transform these successful solutions into learning experiences that the model can practice and learn from under a general harness. Trajectory rewriting involves three distinct model roles: a planner, a critic, and an executor.

  • Planner: The planner reconstructs a runbook for each successful source trajectory, providing a structured description of how the task was solved. Before planning, the source trajectory is compacted by retaining the task instruction, the model’s actions, and the environment’s observations, while removing harness-specific control messages. The runbook summarizes the required end state, key milestones, useful checks, recovery strategies, and common pitfalls. To reduce direct answer transfer, runbooks describe the task, relevant interfaces, and validation procedures without directly providing the finished deliverable. Multiple runbook candidates are sampled for each source trajectory.
  • Critic: The model itself acts as a critic to filter candidate runbooks before they are used for execution. Deterministic checks are first applied, such as schema validation, removal of known artifacts, and rejection of unsupported tool references. Then, a model-based critic that sees only the public task instruction and the candidate runbook determines whether the runbook provides useful procedure or leaks information that the executor should not receive. Only runbooks that pass this screening stage are retained for rewriting.
  • Executor: For each approved runbook, the executor re-solves the task in a fresh sandbox under the general harness. The runbook is provided as private guidance during generation but is never written into the public trajectory. The executor must therefore produce a new trajectory based on the current environment rather than replaying the source trajectory.

Trajectory Filtering and Verification The model itself is used to flag values that appear in a trajectory but cannot be derived from the task or the environment, which indicates hidden answer transfer. Demonstrations containing such values are discarded. For finetuning, only the public interaction history, comprising the task instruction, environment observations, and the model’s responses, is kept, while the runbook and critic conversation are removed. At inference time, the finetuned model runs under the general-purpose harness alone, without the source harnesses or private runbooks, ensuring that any planning, checking, recovery, or continuation comes from the model itself.

Case Study Examples The authors examine rewritten tasks to demonstrate how a source experience can be transformed into new trajectories under the general harness. As shown in the figure below, the examples illustrate a passing and a failing execution guided by the same runbook in each case, alongside representative command excerpts from the source and both rewrites.

Experiment

The experiments evaluate a self-improvement pipeline that collects successful terminal-task rollouts from three complementary execution harnesses, Terminus 2, StateM, and Recursive Self-Reflect Terminus, then reconstructs those experiences into standardized trajectories under the general harness. They show that different harnesses unlock different problem-solving behaviors and jointly expand task coverage, while direct fine-tuning on source rollouts produces mixed gains and can introduce looping failures. Rewriting successful trajectories through recursive self-rewrite leads to cleaner, more generalizable training data and improves performance over both the base model and direct fine-tuning, with additional partial progress on long-horizon tasks.

Combining rollouts from Terminus 2, Recursive Self-Reflect Terminus, and StateM expands solved-task coverage beyond any single harness or pair. The three-harness union solves the most tasks overall and adds a substantial number over the strongest individual harness, with relative gains on both RST and SWR. Individual harnesses show complementary strengths, as leaders vary across benchmark splits. The three-harness union yields the highest overall task coverage and a roughly one-third relative increase over the strongest individual harness. Across individual harnesses, Recursive Self-Reflect Terminus leads on RST, while Terminus 2 leads on SWR and the pooled set; pairwise unions consistently exceed any single harness.

Combining rollouts from Terminus 2, Recursive Self-Reflect Terminus, and StateM broadens discovery coverage beyond any single harness. The pooled collection contains 2,001 successful trajectories covering 759 distinct tasks, which is more solved tasks than the strongest individual harness. Per-rollout success rates are similar across harnesses, indicating that multi-harness pooling mainly increases task coverage. The strongest individual harness solves 565 tasks, while the pooled union solves 759 tasks, adding 194 solved tasks and a relative coverage increase of 34.3%. Recursive Self-Reflect Terminus achieves the highest individual rollout success rate at 15.9%, StateM has the lowest at 11.7%, and the pooled set reaches 13.7%.

Across passing trajectories, the three harnesses show distinct execution styles. Terminus 2 produces shorter trajectories with more exploration-oriented commands and more passing rollouts, while StateM produces longer trajectories with more commands per turn but less exploration. RSRT sits between them on trajectory length and exploration, with the most completion claims per trajectory and a small share of passes occurring after rejection. Terminus 2 has the largest number of passing trajectories and the shortest average and median turn counts among the three harnesses. StateM shows the longest median trajectory length and highest commands per turn, but the lowest exploration command share. RSRT records the most completion claims per trajectory and is the only harness with a notable percentage of passes after rejection.

Rewrites guided by the same runbook tend to have more similar command-level behavior than rewrites using different runbooks. This pattern is stronger for Markdown, while OpenFOAM shows a weaker effect for command metrics. Workflow-level similarity is less consistent and varies by task. Same-runbook pairs show higher exact command overlap and command-order similarity than different-runbook pairs, with a clearer gap for Markdown. Tool-transition and action-sequence similarity are not consistently higher within runbooks; OpenFOAM shows little tool-transition difference and slightly lower within-runbook action-sequence similarity.

RSR outperforms both the base model and Direct SFT across the reported benchmarks, with higher pass@3, higher mean per-run pass rate, and better process reward on LHTB. Direct SFT shows mixed results, improving on several benchmarks but declining on TB2, where training without rewriting can introduce looping behaviors. The results suggest that scaling and standardizing successful experiences under a general harness strengthens generalization. RSR achieves the highest pass@3 and mean per-run pass rate across all five reported benchmarks. Direct SFT has mixed results, improving over the base model on TBH, TB3, and TB4 but declining on TB2 and falling below its base mean per-run pass rate.

The experiments evaluate multi-harness rollouts from Terminus 2, Recursive Self-Reflect Terminus, and StateM, along with runbook-guided rewrite similarity and a comparison of RSR against base and Direct SFT models. Combining the three harnesses yields complementary task coverage and the largest set of solved tasks, while the harnesses show distinct execution styles and same-runbook rewrites are more similar at the command level, especially for Markdown. RSR consistently outperforms the base model and Direct SFT, whereas Direct SFT improves on several benchmarks but can degrade on others through looping behaviors.


AIでAIを構築

アイデアからローンチまで — 無料のAIコーディング支援、すぐに使える環境、最高のGPU価格でAI開発を加速。

AI コーディング補助
すぐに使える GPU
最適な料金体系

HyperAI Newsletters

最新情報を購読する
北京時間 毎週月曜日の午前9時 に、その週の最新情報をメールでお届けします
メール配信サービスは MailChimp によって提供されています