HyperAIHyperAI

Command Palette

Search for a command to run...

AREX-2 : faire progresser les agents auto-améliorants grâce à des tâches réflexives à long horizon

Résumé

Nous présentons AREX-2, un travail visant à renforcer la capacité d'auto-amélioration des agents LLM, que nous définissons comme l'aptitude à raffiner itérativement une solution au moment du test. Cette capacité repose sur deux aptitudes complémentaires : la réflexion, qui produit une solution meilleure que la solution actuelle, et l'exécution à long horizon, qui maintient l'efficacité de l'itération sur de nombreux tours. Nous faisons l'hypothèse que ces deux aptitudes sont indépendantes du domaine et peuvent donc être apprises dans des scénarios bien adaptés à la supervision. En conséquence, nous synthétisons des trajectoires d'amélioration à long horizon à partir de tâches d'apprentissage automatique et de programmation algorithmique, deux domaines qui offrent une rétroaction vérifiable et récompensent une itération soutenue. Entraîné sur ces données, notre agent, construit sur Qwen3.8-27B, obtient de solides résultats sur MLE-bench Lite (81,8) et Frontier-CS (70,7), transfère vers la recherche approfondie avec 84,0 sur BrowseComp, 52,6 sur HLE, 92,2 sur GAIA et 93,8 sur DeepSearchQA, et continue de s'améliorer à mesure que son budget de tours augmente. Ces résultats montrent que les données réflexives à long horizon constituent une voie efficace vers des agents auto-améliorants.

One-sentence Summary

The AREX Team at Beijing Academy of Artificial Intelligence (BAAI) presents AREX-2, a self-improving LLM agent built on Qwen3.8-27B and trained on synthesized long-horizon reflective trajectories from machine learning and algorithmic programming, which achieves 81.8 on MLE-bench Lite and 70.7 on Frontier-CS while transferring to deep research with 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA.

Key Contributions

  • The paper introduces AREX-2, a self-improving agent built on Qwen3.8-27B that is trained on synthesized long-horizon improvement trajectories from machine learning engineering and algorithmic programming, with trajectories selected by final outcome so failed runs and regressions remain in the data.
  • The method reaches 81.8 on MLE-bench Lite and 70.7 on Frontier-CS, with Frontier-CS reported as the highest among open-weight models compared.
  • Without added deep-research training data, AREX-2 transfers to deep research and reaches 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA, and its performance keeps improving as the round budget grows; a stage-wise ablation attributes gains to operational knowledge, training, and a larger budget.

Introduction

Hard problems are solved through iterative loops of trying, measuring, and revising, but in most systems that loop is implemented in the scaffolding rather than learned by the model. Current training data for agents is mostly built one attempt at a time, discarding intermediate attempts, feedback, and revisions, so models learn what correct solutions look like but not how to improve them over many rounds. The authors define self-improvement as the ability to turn more rounds on a task into a better solution by the model's own judgment, and they decompose it into reflection and long-horizon execution. To supervise these capabilities, they construct long trajectories in machine learning engineering and algorithmic programming, keeping whole trajectories selected by final outcome rather than individual step success. They train AREX-2 from Qwen3.8-27B on this data, reaching 81.8 on MLE-bench Lite and 70.7 on Frontier-CS, and show that the learned long-horizon reflection transfers to deep research without new training data.

Dataset

The authors build environments from two source types:

  • GitHub repositories: supply machine learning tasks. A teacher model reads the repository, identifies the metric it measures, and writes a task to improve that metric. It also creates a scoring script that computes the metric on a held-out split the agent cannot access, plus a sandbox with the repository installed and its data in place.
  • Online judges: supply algorithmic programming tasks. The task is the problem statement, the sandbox contains a compiler and sample cases, and the score comes from hidden tests. Where allowed, the score is graded rather than binary, for example the fraction of tests passed or the quality of a heuristic solution relative to the best known one.

Each source is converted into an environment only if it passes two execution-based checks:

  • The reference solution, either the repository’s own model or the judge’s accepted submission, must run under the scoring setup and obtain a score.
  • A simple baseline must score well below the reference, showing that the environment leaves room for improvement.
  • Environments that fail this second condition are dropped because they cannot produce trajectories of sustained improvement.

From each admitted environment, the authors construct trajectories by running a teacher agent under conditions designed for long-horizon reflection:

  • More rounds: the agent receives a large round and wall-clock budget, on the order of hours and hundreds of tool calls per task, and is told what the budget is.
  • Operational knowledge: the agent learns how the environment works by reading source and documentation, searching for related papers and libraries, and running small experiments. Some knowledge is supplied as skills: general skills prepared in advance from machine learning repositories, and task-specific skills built by searching for related information while excluding material about the task itself.
  • Feedback in the loop: each round ends with a submission, and the agent sees the score and environment report before deciding its next step. Rounds that lower the score are kept, because the agent’s response to failure is part of what the model should learn.

The resulting trajectory data consists of many rounds of reasoning, searches, failures, recoveries, and submissions recorded while a capable agent works on one task. In training, the model learns from these trajectories how to investigate an environment, use skills, and convert a larger round budget into cumulative improvement. Skills remain available at test time, and the round budget is the same resource that is scaled at test time.

The provided excerpt does not specify exact dataset sizes, training split ratios, mixture ratios, or cropping strategy.

Method

The authors formalize self-improvement as a test-time scaling problem. An environment is defined as a pair e=(σ,S)e=(\sigma, S)e=(σ,S), where σ\sigmaσ is a task and SSS is a scoring function that maps a candidate solution yyy to a scalar score S(y)S(y)S(y). The agent works over multiple rounds. In round ttt, it submits a solution yty_tyt​ and receives feedback ftf_tft​, which includes the score and any environment reports such as logs, errors, or timings. A trajectory of nnn rounds is

τ=(y0,f0,y1,f1,…,yn,fn),\tau = (y_0, f_0, y_1, f_1, \dots, y_n, f_n),τ=(y0​,f0​,y1​,f1​,…,yn​,fn​),

and the best score reached after ttt rounds is

st=max⁡k≤tS(yk).s_t = \max_{k \le t} S(y_k).st​=k≤tmax​S(yk​).

Under a policy πθ\pi_\thetaπθ​, the expected gain of round ttt is

rt=Eπθ[st−st−1].r_t = \mathbb{E}_{\pi_\theta}[s_t - s_{t-1}].rt​=Eπθ​​[st​−st−1​].

For a fixed initial score s0s_0s0​ and a test-time budget of TTT rounds, the expected total improvement is the sum of these per-round gains:

Eπθ[sT]−s0=∑t=1Trt.\mathbb{E}_{\pi_\theta}[s_T] - s_0 = \sum_{t=1}^{T} r_t.Eπθ​​[sT​]−s0​=t=1∑T​rt​.

The budget TTT is therefore the resource that is scaled at test time. The policy controls how much of that budget can be used productively. For a small threshold ϵ\epsilonϵ, the authors define the number of productive rounds and the mean gain over those rounds as

T∗=∣{t≤T:rt>ϵ}∣,rˉ=1T∗∑t≤T:rt>ϵrt.T^* = \left|\{t \le T : r_t > \epsilon\}\right|, \quad \bar{r} = \frac{1}{T^*} \sum_{t \le T : r_t > \epsilon} r_t.T∗=∣{t≤T:rt​>ϵ}∣,rˉ=T∗1​t≤T:rt​>ϵ∑​rt​.

The total improvement is approximately rˉT∗\bar{r} T^*rˉT∗. Reflection determines the average gain rˉ\bar{r}rˉ, while long-horizon execution determines T∗T^*T∗. Thus, a larger budget helps only up to the point where additional rounds stop producing meaningful gains.

The central hypothesis is that long-horizon reflection is a meta-skill. The behaviors that produce sustained improvement, using feedback, acquiring operational knowledge, and recovering from setbacks, transfer across environments even when the tasks and scoring functions differ. The authors train the policy by imitation on trajectories that exhibit these behaviors.

Environment construction starts from two sources: GitHub repositories for machine learning tasks and online judges for algorithmic programming tasks. A teacher model converts each source into an environment (σ,S)(\sigma, S)(σ,S). For a repository, the teacher reads the code, identifies the metric being measured, writes a task that asks the agent to improve that metric, and creates a scoring script that computes the metric on a held-out split. It also builds a sandbox with the repository installed and its data in place. For a judge problem, the task is the problem statement, the sandbox contains a compiler and sample cases, and the score is the judge's hidden-test result. Where possible, the judge score is graded rather than binary, such as the fraction of tests passed or the quality of a heuristic solution relative to the best known one, so that each round is informative.

An environment is admitted only if two execution-based checks pass. First, the reference solution that comes with the source must run under the scoring function and obtain a valid score. Second, a simple baseline must score well below the reference, ensuring that the environment leaves room for improvement. Environments that fail the second check are dropped because they cannot produce trajectories of sustained improvement.

To construct trajectories, the authors run a teacher agent under three conditions designed to elicit long-horizon reflection. The first is a large round and wall-clock budget, on the order of hours and hundreds of tool calls per task. The agent is told the budget, which allows it to establish a baseline, use early rounds for measurement, and use later rounds for revisions. The second condition is operational knowledge acquisition. The agent must learn how the environment works by reading source code and documentation, searching for relevant papers and libraries, and running small experiments. Some operational knowledge is also supplied through compact skill documents placed in the agent's context. General skills are prepared from machine learning repositories and common practices, while task-specific skills are built from search results related to the task, excluding material about the task itself. The third condition is a feedback loop. Each round ends with a submission, and the agent sees the score and environment report before deciding what to do next. This is where reflection occurs: the agent compares the observed result with its expectation, keeps or reverts the change, and forms the next hypothesis. Rounds that lower the score are not removed, because the response to failure is precisely what the model should learn.

For training, trajectories are selected as whole units. A trajectory is kept if its final score is high enough relative to the reference and if its process is well formed, following the user-assistant-tool format, with every tool call followed by an observation, proper termination, and no violation of task rules. No condition is placed on individual rounds, so kept trajectories may contain failed runs, regressions, and abandoned approaches.

During supervision, the whole trajectory remains as context, and the loss is applied only to decisions that move the solution forward. Outputs are grouped into steps, where each step is one decision together with the actions issued before the next observation. The loss is applied to decisions that diagnose failures, repair them, change strategy, run the next experiment, or submit an improved solution. Steps that make no progress, repeated polling, calls with no new observation, near-duplicate turns, system messages, environment reports, and retrieved documents receive no loss. The model therefore learns how to respond to a setback and turn it into later improvement while the failure itself remains visible in context. The selected trajectories are combined with the previous deep-research data, and Qwen3.8-27B is fine-tuned on the mixture to obtain the final model.

Experiment

AREX-2 is evaluated across six benchmarks covering algorithmic programming, machine learning engineering, deep research, and general agentic reasoning. The overall results show strong performance in the source training domains and clear cross-domain transfer to deep research tasks without new research-specific training data, supporting the view that long-horizon reflection is a transferable meta-skill. Follow-up scaling experiments indicate that the agent continues to improve with larger compute budgets both with and without external correctness feedback, while ablations confirm that operational skills, model training, and additional rounds each contribute meaningfully to performance.

AREX-2 achieves the highest MLE-Lite Any Medal rate among compared systems, outperforming the strongest open-weight baseline and the strongest closed-weight model by clear margins. It also performs competitively on Frontier-CS, exceeding open-weight alternatives and coming close to the best closed-weight system. At 27B parameters, it combines leading machine learning engineering results with strong algorithmic coding performance. AREX-2 leads all compared systems on MLE-Lite, surpassing both the strongest open-weight and strongest closed-weight baselines. On Frontier-CS, AREX-2 exceeds the strongest open-weight baseline and approaches the top closed-weight model while using 27B parameters.

AREX-2 achieves the strongest reported small-model results on most deep research benchmarks, while remaining competitive with much larger systems. It also surpasses both earlier AREX models on all four benchmarks, despite using unchanged deep-research training data. Frontier-scale models still hold higher scores on BrowseComp and HLE. AREX-2 leads all reported models at most 40B parameters on BrowseComp, text-only HLE, and DeepSearchQA. On GAIA, it ranks behind XYZ-Aquila-mini and Agents-A1 but above the remaining small models with reported results. It outperforms both previous recipe AREX models on all four benchmarks while using the same deep-research training data. It is competitive with larger systems, exceeding DeepSeek-V4-Pro and Kimi-K2.6 on BrowseComp and ranking third on DeepSearchQA.

The experiments evaluate AREX-2, a 27B-parameter model, on machine learning engineering and algorithmic coding tasks as well as deep research benchmarks. In the first setting, AREX-2 achieves the highest MLE-Lite Any Medal rate among all compared systems and outperforms strong open-weight and closed-weight baselines, while remaining competitive on Frontier-CS. In deep research benchmarks, it leads most reported models at or below 40B parameters on BrowseComp, text-only HLE, and DeepSearchQA, surpasses both earlier AREX models on all four benchmarks despite unchanged training data, and stays competitive with larger systems, though frontier-scale models still score higher on BrowseComp and HLE.


Créer de l'IA avec l'IA

De l'idée au lancement — accélérez votre développement IA avec le co-codage IA gratuit, un environnement prêt à l'emploi et le meilleur prix pour les GPU.

Codage assisté par IA
GPU prêts à l’emploi
Tarifs les plus avantageux

HyperAI Newsletters

Abonnez-vous à nos dernières mises à jour
Nous vous enverrons les dernières mises à jour de la semaine dans votre boîte de réception à neuf heures chaque lundi matin
Propulsé par MailChimp