HyperAIHyperAI

Command Palette

Search for a command to run...

AREX-2: تطوير الوكلاء ذاتيي التحسين عبر مهام تأملية طويلة الأفق

الملخص

نقدم AREX-2، وهو جهد لتطوير قدرة التحسين الذاتي لدى وكلاء النماذج اللغوية الكبيرة، والتي نعرّفها بأنها القدرة على تنقيح الحل تكرارياً في وقت الاختبار. تستند هذه القدرة إلى قدرتين متكاملتين: التأمل الذي ينتج حلاً أفضل من الحل الحالي، والتنفيذ طويل الأفق الذي يُبقي التكرار فعالاً عبر جولات عديدة. نفترض أن كلتا القدرتين غير مرتبطتين بمجال محدد، وبالتالي يمكن تعلمهما في سياقات مناسبة للإشراف. بناءً على ذلك، نولّد مسارات تحسين طويلة الأفق من مهام تعلم الآلة والبرمجة الخوارزمية، وهما مجالان يوفران تغذية راجعة قابلة للتحقق ويكافئان التكرار المستمر. بعد التدريب على هذه البيانات، يحقق وكيلنا المبني على Qwen3.8-27B نتائج قوية على MLE-bench Lite (81.8) وFrontier-CS (70.7)، وينتقل إلى البحث العميق محققاً 84.0 على BrowseComp و52.6 على HLE و92.2 على GAIA و93.8 على DeepSearchQA، ويستمر في التحسن كلما زادت ميزانية الجولات. تُظهر هذه النتائج أن البيانات التأملية طويلة الأفق طريق فعال نحو وكلاء ذاتيي التحسين.

One-sentence Summary

The AREX Team at Beijing Academy of Artificial Intelligence (BAAI) presents AREX-2, a self-improving LLM agent built on Qwen3.8-27B and trained on synthesized long-horizon reflective trajectories from machine learning and algorithmic programming, which achieves 81.8 on MLE-bench Lite and 70.7 on Frontier-CS while transferring to deep research with 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA.

Key Contributions

  • The paper introduces AREX-2, a self-improving agent built on Qwen3.8-27B that is trained on synthesized long-horizon improvement trajectories from machine learning engineering and algorithmic programming, with trajectories selected by final outcome so failed runs and regressions remain in the data.
  • The method reaches 81.8 on MLE-bench Lite and 70.7 on Frontier-CS, with Frontier-CS reported as the highest among open-weight models compared.
  • Without added deep-research training data, AREX-2 transfers to deep research and reaches 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA, and its performance keeps improving as the round budget grows; a stage-wise ablation attributes gains to operational knowledge, training, and a larger budget.

Introduction

Hard problems are solved through iterative loops of trying, measuring, and revising, but in most systems that loop is implemented in the scaffolding rather than learned by the model. Current training data for agents is mostly built one attempt at a time, discarding intermediate attempts, feedback, and revisions, so models learn what correct solutions look like but not how to improve them over many rounds. The authors define self-improvement as the ability to turn more rounds on a task into a better solution by the model's own judgment, and they decompose it into reflection and long-horizon execution. To supervise these capabilities, they construct long trajectories in machine learning engineering and algorithmic programming, keeping whole trajectories selected by final outcome rather than individual step success. They train AREX-2 from Qwen3.8-27B on this data, reaching 81.8 on MLE-bench Lite and 70.7 on Frontier-CS, and show that the learned long-horizon reflection transfers to deep research without new training data.

Dataset

The authors build environments from two source types:

  • GitHub repositories: supply machine learning tasks. A teacher model reads the repository, identifies the metric it measures, and writes a task to improve that metric. It also creates a scoring script that computes the metric on a held-out split the agent cannot access, plus a sandbox with the repository installed and its data in place.
  • Online judges: supply algorithmic programming tasks. The task is the problem statement, the sandbox contains a compiler and sample cases, and the score comes from hidden tests. Where allowed, the score is graded rather than binary, for example the fraction of tests passed or the quality of a heuristic solution relative to the best known one.

Each source is converted into an environment only if it passes two execution-based checks:

  • The reference solution, either the repository’s own model or the judge’s accepted submission, must run under the scoring setup and obtain a score.
  • A simple baseline must score well below the reference, showing that the environment leaves room for improvement.
  • Environments that fail this second condition are dropped because they cannot produce trajectories of sustained improvement.

From each admitted environment, the authors construct trajectories by running a teacher agent under conditions designed for long-horizon reflection:

  • More rounds: the agent receives a large round and wall-clock budget, on the order of hours and hundreds of tool calls per task, and is told what the budget is.
  • Operational knowledge: the agent learns how the environment works by reading source and documentation, searching for related papers and libraries, and running small experiments. Some knowledge is supplied as skills: general skills prepared in advance from machine learning repositories, and task-specific skills built by searching for related information while excluding material about the task itself.
  • Feedback in the loop: each round ends with a submission, and the agent sees the score and environment report before deciding its next step. Rounds that lower the score are kept, because the agent’s response to failure is part of what the model should learn.

The resulting trajectory data consists of many rounds of reasoning, searches, failures, recoveries, and submissions recorded while a capable agent works on one task. In training, the model learns from these trajectories how to investigate an environment, use skills, and convert a larger round budget into cumulative improvement. Skills remain available at test time, and the round budget is the same resource that is scaled at test time.

The provided excerpt does not specify exact dataset sizes, training split ratios, mixture ratios, or cropping strategy.

Method

The authors formalize self-improvement as a test-time scaling problem. An environment is defined as a pair e=(σ,S)e=(\sigma, S)e=(σ,S), where σ\sigmaσ is a task and SSS is a scoring function that maps a candidate solution yyy to a scalar score S(y)S(y)S(y). The agent works over multiple rounds. In round ttt, it submits a solution yty_tyt​ and receives feedback ftf_tft​, which includes the score and any environment reports such as logs, errors, or timings. A trajectory of nnn rounds is

τ=(y0,f0,y1,f1,…,yn,fn),\tau = (y_0, f_0, y_1, f_1, \dots, y_n, f_n),τ=(y0​,f0​,y1​,f1​,…,yn​,fn​),

and the best score reached after ttt rounds is

st=max⁡k≤tS(yk).s_t = \max_{k \le t} S(y_k).st​=k≤tmax​S(yk​).

Under a policy πθ\pi_\thetaπθ​, the expected gain of round ttt is

rt=Eπθ[st−st−1].r_t = \mathbb{E}_{\pi_\theta}[s_t - s_{t-1}].rt​=Eπθ​​[st​−st−1​].

For a fixed initial score s0s_0s0​ and a test-time budget of TTT rounds, the expected total improvement is the sum of these per-round gains:

Eπθ[sT]−s0=∑t=1Trt.\mathbb{E}_{\pi_\theta}[s_T] - s_0 = \sum_{t=1}^{T} r_t.Eπθ​​[sT​]−s0​=t=1∑T​rt​.

The budget TTT is therefore the resource that is scaled at test time. The policy controls how much of that budget can be used productively. For a small threshold ϵ\epsilonϵ, the authors define the number of productive rounds and the mean gain over those rounds as

T∗=∣{t≤T:rt>ϵ}∣,rˉ=1T∗∑t≤T:rt>ϵrt.T^* = \left|\{t \le T : r_t > \epsilon\}\right|, \quad \bar{r} = \frac{1}{T^*} \sum_{t \le T : r_t > \epsilon} r_t.T∗=∣{t≤T:rt​>ϵ}∣,rˉ=T∗1​t≤T:rt​>ϵ∑​rt​.

The total improvement is approximately rˉT∗\bar{r} T^*rˉT∗. Reflection determines the average gain rˉ\bar{r}rˉ, while long-horizon execution determines T∗T^*T∗. Thus, a larger budget helps only up to the point where additional rounds stop producing meaningful gains.

The central hypothesis is that long-horizon reflection is a meta-skill. The behaviors that produce sustained improvement, using feedback, acquiring operational knowledge, and recovering from setbacks, transfer across environments even when the tasks and scoring functions differ. The authors train the policy by imitation on trajectories that exhibit these behaviors.

Environment construction starts from two sources: GitHub repositories for machine learning tasks and online judges for algorithmic programming tasks. A teacher model converts each source into an environment (σ,S)(\sigma, S)(σ,S). For a repository, the teacher reads the code, identifies the metric being measured, writes a task that asks the agent to improve that metric, and creates a scoring script that computes the metric on a held-out split. It also builds a sandbox with the repository installed and its data in place. For a judge problem, the task is the problem statement, the sandbox contains a compiler and sample cases, and the score is the judge's hidden-test result. Where possible, the judge score is graded rather than binary, such as the fraction of tests passed or the quality of a heuristic solution relative to the best known one, so that each round is informative.

An environment is admitted only if two execution-based checks pass. First, the reference solution that comes with the source must run under the scoring function and obtain a valid score. Second, a simple baseline must score well below the reference, ensuring that the environment leaves room for improvement. Environments that fail the second check are dropped because they cannot produce trajectories of sustained improvement.

To construct trajectories, the authors run a teacher agent under three conditions designed to elicit long-horizon reflection. The first is a large round and wall-clock budget, on the order of hours and hundreds of tool calls per task. The agent is told the budget, which allows it to establish a baseline, use early rounds for measurement, and use later rounds for revisions. The second condition is operational knowledge acquisition. The agent must learn how the environment works by reading source code and documentation, searching for relevant papers and libraries, and running small experiments. Some operational knowledge is also supplied through compact skill documents placed in the agent's context. General skills are prepared from machine learning repositories and common practices, while task-specific skills are built from search results related to the task, excluding material about the task itself. The third condition is a feedback loop. Each round ends with a submission, and the agent sees the score and environment report before deciding what to do next. This is where reflection occurs: the agent compares the observed result with its expectation, keeps or reverts the change, and forms the next hypothesis. Rounds that lower the score are not removed, because the response to failure is precisely what the model should learn.

For training, trajectories are selected as whole units. A trajectory is kept if its final score is high enough relative to the reference and if its process is well formed, following the user-assistant-tool format, with every tool call followed by an observation, proper termination, and no violation of task rules. No condition is placed on individual rounds, so kept trajectories may contain failed runs, regressions, and abandoned approaches.

During supervision, the whole trajectory remains as context, and the loss is applied only to decisions that move the solution forward. Outputs are grouped into steps, where each step is one decision together with the actions issued before the next observation. The loss is applied to decisions that diagnose failures, repair them, change strategy, run the next experiment, or submit an improved solution. Steps that make no progress, repeated polling, calls with no new observation, near-duplicate turns, system messages, environment reports, and retrieved documents receive no loss. The model therefore learns how to respond to a setback and turn it into later improvement while the failure itself remains visible in context. The selected trajectories are combined with the previous deep-research data, and Qwen3.8-27B is fine-tuned on the mixture to obtain the final model.

Experiment

AREX-2 is evaluated across six benchmarks covering algorithmic programming, machine learning engineering, deep research, and general agentic reasoning. The overall results show strong performance in the source training domains and clear cross-domain transfer to deep research tasks without new research-specific training data, supporting the view that long-horizon reflection is a transferable meta-skill. Follow-up scaling experiments indicate that the agent continues to improve with larger compute budgets both with and without external correctness feedback, while ablations confirm that operational skills, model training, and additional rounds each contribute meaningfully to performance.

AREX-2 achieves the highest MLE-Lite Any Medal rate among compared systems, outperforming the strongest open-weight baseline and the strongest closed-weight model by clear margins. It also performs competitively on Frontier-CS, exceeding open-weight alternatives and coming close to the best closed-weight system. At 27B parameters, it combines leading machine learning engineering results with strong algorithmic coding performance. AREX-2 leads all compared systems on MLE-Lite, surpassing both the strongest open-weight and strongest closed-weight baselines. On Frontier-CS, AREX-2 exceeds the strongest open-weight baseline and approaches the top closed-weight model while using 27B parameters.

AREX-2 achieves the strongest reported small-model results on most deep research benchmarks, while remaining competitive with much larger systems. It also surpasses both earlier AREX models on all four benchmarks, despite using unchanged deep-research training data. Frontier-scale models still hold higher scores on BrowseComp and HLE. AREX-2 leads all reported models at most 40B parameters on BrowseComp, text-only HLE, and DeepSearchQA. On GAIA, it ranks behind XYZ-Aquila-mini and Agents-A1 but above the remaining small models with reported results. It outperforms both previous recipe AREX models on all four benchmarks while using the same deep-research training data. It is competitive with larger systems, exceeding DeepSeek-V4-Pro and Kimi-K2.6 on BrowseComp and ranking third on DeepSearchQA.

The experiments evaluate AREX-2, a 27B-parameter model, on machine learning engineering and algorithmic coding tasks as well as deep research benchmarks. In the first setting, AREX-2 achieves the highest MLE-Lite Any Medal rate among all compared systems and outperforms strong open-weight and closed-weight baselines, while remaining competitive on Frontier-CS. In deep research benchmarks, it leads most reported models at or below 40B parameters on BrowseComp, text-only HLE, and DeepSearchQA, surpasses both earlier AREX models on all four benchmarks despite unchanged training data, and stays competitive with larger systems, though frontier-scale models still score higher on BrowseComp and HLE.


بناء الذكاء الاصطناعي بالذكاء الاصطناعي

من الفكرة إلى الإطلاق — سرّع تطوير الذكاء الاصطناعي الخاص بك مع المساعدة البرمجية المجانية بالذكاء الاصطناعي، وبيئة جاهزة للاستخدام، وأفضل أسعار لوحدات معالجة الرسومات.

البرمجة التعاونية باستخدام الذكاء الاصطناعي
وحدات GPU جاهزة للعمل
أفضل الأسعار

HyperAI Newsletters

اشترك في آخر تحديثاتنا
سنرسل لك أحدث التحديثات الأسبوعية إلى بريدك الإلكتروني في الساعة التاسعة من صباح كل يوم اثنين
مدعوم بواسطة MailChimp