HyperAIHyperAI

Command Palette

Search for a command to run...

コードエージェントRLのためのグループ単位エージェント型評価とアドバンテージ再分配

Jinhao Dong Liang Zhao Zihao Yue Wenhan Ma Linghao Zhang Lei Li Shicheng Li Yifan Song Bowen Ye Fuli Luo

概要

コードエージェントのための強化学習(RL)では、実行可能なテストを用いて二値の報酬が与えられることが多い。このような報酬のもとで、Group Relative Policy Optimization(GRPO)は、各ロールアウトグループ内のテスト合格軌道に同一のアドバンテージを割り当てるため、実装品質やタスク要件への適合性の差が見落とされる。これにより、不必要または範囲外の変更を含む実装よりも、クリーンで的を絞った実装を優先するための学習信号がポリシーに与えられない。我々は、コードエージェントRLにおける品質を考慮したクレジット再分配のためのフレームワークであるGagarを導入する。Gagarは、合格軌道と不合格軌道の両方を含むグループを保持する動的サンプリングに基づいて構築されており、各グループの全軌道を共有ワークスペースに配置し、SFTで訓練されたエージェント型評価器がそれらをまとめて検査してテスト合格候補を順位付けする。この順位に基づき、下位に順位付けされた軌道を低重み化し、全テスト合格軌道のアドバンテージを比例的に再スケーリングすることで元の合計を復元する。この合計保存型の再分配は、品質に基づく低重み化によって確立された相対的重みを保持しつつ、より高品質な実装へクレジットを移行させる。我々は、MiMo-V2.6-Flash(総パラメータ数310B)およびMiMo-V2.6-Pro(総パラメータ数1.02T)のRL実施前SFTチェックポイントを用いて、産業規模でGagarを評価する。コードのみのFlashによる対照実験では、コードエージェント性能の向上、軌道長増加の抑制、およびより安定した訓練が示される。さらに、FlashとProの両方を用いた大規模な混合タスクRLにGagarを適用する。我々の結果は、テストに基づく検証とグループ単位のエージェント型評価を組み合わせることで、コードエージェントRLの品質と安定性が向上することを支持する。

One-sentence Summary

Researchers from Xiaomi, Renmin University of China, Peking University, and other institutions propose Gagar, a quality-aware credit redistribution framework for code agent RL that uses dynamic sampling and an SFT-trained agentic grader to rank test-passing trajectories, downweights lower-ranked trajectories, and applies sum-preserving advantage rescaling; industrial-scale experiments with MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters) show improved code agent performance, reduced trajectory-length growth, and more stable training.

Key Contributions

  • Gagar is a quality-aware credit redistribution framework for code agent RL that uses dynamic sampling to retain groups with both passing and failing trajectories, then uses an SFT-trained agentic grader to inspect and rank test-passing trajectories in a shared workspace.
  • Gagar downweights lower-ranked passing trajectories and proportionally rescales the advantages of all passing trajectories, preserving the original total positive advantage while shifting credit toward higher-quality implementations.
  • Controlled code-only experiments with MiMo-V2.6-Flash show improved code agent performance, reduced trajectory-length growth, and more stable training; mixed-task RL with MiMo-V2.6-Flash and MiMo-V2.6-Pro reaches DeepSWE v1.1 avg@3 scores of 67.9 and 71.9.

Introduction

Reinforcement learning with executable feedback is a scalable way to train code agents that inspect repositories, modify code, and validate changes over long interactions. In common group-relative setups such as GRPO, all test-passing trajectories receive the same positive outcome advantage, so the training signal cannot distinguish cleaner, more focused patches from overly complex or risky implementations. Static text-based assessment also tends to miss repository context and execution evidence. The authors introduce Gagar, a quality-aware framework that uses an agentic grader to inspect code, run targeted checks, and rank test-passing candidates within each rollout group, then redistributes credit among successful trajectories while preserving their total positive advantage. This approach is designed to reinforce precise, minimally invasive, merge-ready solutions and is validated at industrial scale with MiMo models.

Method

The authors leverage groupwise quality supervision to augment reinforcement learning for code agents, moving beyond binary task outcomes. As shown in the figure below, the framework integrates groupwise agentic grading into the RL training loop to evaluate and rank valid passing implementations.

In the training setup, for each coding task, the rollout policy generates multiple trajectories. The final patch from each trajectory is evaluated using executable tests to yield a binary outcome reward. Let Gx={τi}i=1n\mathcal{G}_x = \{\tau_i\}_{i=1}^nGx​={τi​}i=1n​ denote the group of valid trajectories. The authors use the mean-centered outcome advantage Ai=Ri−RˉA_i = R_i - \bar{R}Ai​=Ri​−Rˉ, where Rˉ=n−1∑j=1nRj\bar{R} = n^{-1} \sum_{j=1}^n R_jRˉ=n−1∑j=1n​Rj​. To ensure both successful and failed trajectories are present, dynamic sampling is adopted so that 0<Rˉ<10 < \bar{R} < 10<Rˉ<1. All passing trajectories initially receive the same positive advantage 1−Rˉ1 - \bar{R}1−Rˉ, which fails to differentiate implementation quality.

To address this, the authors introduce groupwise agentic grading. The grader jointly examines all rollouts in a shared workspace containing the task specification, repository, complete trajectories, submitted patches, and test outputs. This comparison helps identify the strongest solutions and exposes ineffective strategies. The agentic grader gathers evidence iteratively, reviewing turn-by-turn summaries, reading relevant trajectory portions, and cross-checking patches against repository code and test logs. It can also run targeted checks. Quality is assessed across five criteria: approach suitability, precision, minimality, side effects, and codebase consistency. Passing solutions are checked for leaked or external answers, and confirmed hacks receive zero reward. The remaining candidates are scored and mapped to three quality tiers, which determine a discount factor fi∈(0,1]f_i \in (0, 1]fi​∈(0,1] on their positive advantage.

Following grading, sum-preserving advantage redistribution is applied. Simply applying the discount factors would yield A~i=fiAi\tilde{A}_i = f_i A_iA~i​=fi​Ai​ for passing trajectories, removing positive credit without a corresponding change on the negative side. This deficit weakens the reinforcement of successful trajectories. To preserve the total positive advantage S+=∑i∈PAiS_+ = \sum_{i \in \mathcal{P}} A_iS+​=∑i∈P​Ai​ assigned to passing trajectories, the authors apply a common rescaling factor:

λ=S+∑j∈PfjAj\lambda = \frac{S_+}{\sum_{j \in \mathcal{P}} f_j A_j}λ=∑j∈P​fj​Aj​S+​​

The redistributed advantage is then:

Ai⋆={λfiAi,i∈P,Ai,i∈F.A_i^\star = \begin{cases} \lambda f_i A_i, & i \in \mathcal{P}, \\ A_i, & i \in \mathcal{F}. \end{cases}Ai⋆​={λfi​Ai​,Ai​,​i∈P,i∈F.​

This ensures the total positive advantage is restored and the change is zero-sum over the passing subset, preserving both the ordering and relative strength of quality preferences.

During online training integration, grading runs asynchronously with rollout collection. The system checks the validity of grading results, falling back to original outcome advantages if necessary. Confirmed reliance on external solutions resets the affected reward to zero before group statistics are recomputed. The resulting sequence-level advantages are broadcast to the model-generated response tokens to supervise the RL policy update.

Experiment

These experiments evaluate Gagar on code-only reinforcement learning with MiMo-V2.6-Flash and on industrial-scale mixed-task RL with Flash and Pro, using DeepSWE v1.1 and SWE-bench Pro as benchmarks. The main comparisons show that Gagar improves long-horizon code agent performance, keeps trajectories shorter and more efficient, and yields higher implementation quality in external model review. An ablation confirms that sum-preserving redistribution is important for training stability, since downweighting passing trajectories without restoring credit leads to unstable dynamics and degraded downstream behavior. Industrial-scale runs further demonstrate competitive code agent performance relative to larger frontier models.

After industrial-scale mixed-task RL with Gagar, MiMo-V2.6-Flash and MiMo-V2.6-Pro achieve coding benchmark results competitive with strong external baselines. MiMo-V2.6-Pro outperforms Kimi K3 on DeepSWE despite having substantially fewer total parameters and surpasses GPT-5.6 Sol on SWE-bench Pro. Its DeepSWE score approaches GPT-5.6 Sol and Claude Opus 5, while it trails Claude Opus 5 on SWE-bench Pro. MiMo-V2.6-Pro outperforms Kimi K3 on DeepSWE despite using substantially fewer total parameters. MiMo-V2.6-Pro surpasses GPT-5.6 Sol on SWE-bench Pro. MiMo-V2.6-Pro approaches GPT-5.6 Sol and Claude Opus 5 on DeepSWE but remains below Claude Opus 5 on SWE-bench Pro. MiMo-V2.6-Flash achieves DeepSWE and SWE-bench Pro scores close to larger external models.

This experiment evaluates MiMo-V2.6-Flash and MiMo-V2.6-Pro after industrial-scale mixed-task RL with Gagar on coding benchmarks. The results show that MiMo-V2.6-Pro is competitive with strong external models, outperforming Kimi K3 on DeepSWE despite using substantially fewer parameters and surpassing GPT-5.6 Sol on SWE-bench Pro, while approaching GPT-5.6 Sol and Claude Opus 5 on DeepSWE but remaining below Claude Opus 5 on SWE-bench Pro. MiMo-V2.6-Flash also achieves DeepSWE and SWE-bench Pro scores close to larger external models.


AIでAIを構築

アイデアからローンチまで — 無料のAIコーディング支援、すぐに使える環境、最高のGPU価格でAI開発を加速。

AI コーディング補助
すぐに使える GPU
最適な料金体系

HyperAI Newsletters

最新情報を購読する
北京時間 毎週月曜日の午前9時 に、その週の最新情報をメールでお届けします
メール配信サービスは MailChimp によって提供されています