Command Palette
Search for a command to run...
التقييم الوكيل على مستوى المجموعات وإعادة توزيع الميزة في التعلم المعزز لوكلاء الكود
التقييم الوكيل على مستوى المجموعات وإعادة توزيع الميزة في التعلم المعزز لوكلاء الكود
Jinhao Dong Liang Zhao Zihao Yue Wenhan Ma Linghao Zhang Lei Li Shicheng Li Yifan Song Bowen Ye Fuli Luo
الملخص
غالبًا ما يستخدم التعلم المعزز (RL) لوكلاء الكود اختبارات قابلة للتنفيذ لتوفير مكافآت ثنائية. ومع هذه المكافآت، يُسنِد أسلوب تحسين السياسة النسبي المجموعي (GRPO) ميزات متطابقة للمسارات الناجحة في الاختبارات داخل كل مجموعة توليد، متجاهلًا الاختلافات في جودة التنفيذ والالتزام بمتطلبات المهمة. وهذا يترك السياسة دون إشارة تعلم تفضّل التنفيذات النظيفة والمركّزة على تلك التي تحتوي على تغييرات غير ضرورية أو خارج نطاق المهمة. نقدم Gagar، إطارًا لإعادة توزيع الفضل المراعي للجودة في التعلم المعزز لوكلاء الكود. يُبنى Gagar على عينات ديناميكية تحتفظ بالمجموعات التي تحتوي على مسارات ناجحة وفاشلة معًا، ويضع جميع المسارات من كل مجموعة في مساحة عمل مشتركة، حيث يفحصها مُقيّم وكيل مُدرَّب عبر الضبط الدقيق الموجَّه (SFT) بشكل مشترك ويُرتب المرشحين الناجحين في الاختبارات. واستنادًا إلى هذا الترتيب، نخفض وزن المسارات ذات الترتيب الأدنى ونعيد قياس ميزات جميع المسارات الناجحة في الاختبارات على نحو تناسبي لاستعادة مجموعها الأصلي. وتحافظ إعادة التوزيع الحافظة للمجموع على الأوزان النسبية الناتجة عن تخفيض الوزن القائم على الجودة، مع تحويل الفضل نحو التنفيذات الأعلى جودة. نقيّم Gagar على نطاق صناعي باستخدام نقاط فحص SFT قبل التعلم المعزز لنموذجي MiMo-V2.6-Flash (بإجمالي 310 مليار معلمة) وMiMo-V2.6-Pro (بإجمالي 1.02 تريليون معلمة). تُظهر التجارب المضبوطة على Flash والمقتصرة على الكود تحسنًا في أداء وكيل الكود، وانخفاضًا في نمو أطوال المسارات، وتدريبًا أكثر استقرارًا. كما نطبق Gagar في تعلم معزز واسع النطاق ومتعدد المهام باستخدام كل من Flash وPro. تدعم نتائجنا الجمع بين التحقق القائم على الاختبارات والتقييم الوكيل على مستوى المجموعات لتحسين جودة واستقرار التعلم المعزز لوكلاء الكود.
One-sentence Summary
Researchers from Xiaomi, Renmin University of China, Peking University, and other institutions propose Gagar, a quality-aware credit redistribution framework for code agent RL that uses dynamic sampling and an SFT-trained agentic grader to rank test-passing trajectories, downweights lower-ranked trajectories, and applies sum-preserving advantage rescaling; industrial-scale experiments with MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters) show improved code agent performance, reduced trajectory-length growth, and more stable training.
Key Contributions
- Gagar is a quality-aware credit redistribution framework for code agent RL that uses dynamic sampling to retain groups with both passing and failing trajectories, then uses an SFT-trained agentic grader to inspect and rank test-passing trajectories in a shared workspace.
- Gagar downweights lower-ranked passing trajectories and proportionally rescales the advantages of all passing trajectories, preserving the original total positive advantage while shifting credit toward higher-quality implementations.
- Controlled code-only experiments with MiMo-V2.6-Flash show improved code agent performance, reduced trajectory-length growth, and more stable training; mixed-task RL with MiMo-V2.6-Flash and MiMo-V2.6-Pro reaches DeepSWE v1.1 avg@3 scores of 67.9 and 71.9.
Introduction
Reinforcement learning with executable feedback is a scalable way to train code agents that inspect repositories, modify code, and validate changes over long interactions. In common group-relative setups such as GRPO, all test-passing trajectories receive the same positive outcome advantage, so the training signal cannot distinguish cleaner, more focused patches from overly complex or risky implementations. Static text-based assessment also tends to miss repository context and execution evidence. The authors introduce Gagar, a quality-aware framework that uses an agentic grader to inspect code, run targeted checks, and rank test-passing candidates within each rollout group, then redistributes credit among successful trajectories while preserving their total positive advantage. This approach is designed to reinforce precise, minimally invasive, merge-ready solutions and is validated at industrial scale with MiMo models.
Method
The authors leverage groupwise quality supervision to augment reinforcement learning for code agents, moving beyond binary task outcomes. As shown in the figure below, the framework integrates groupwise agentic grading into the RL training loop to evaluate and rank valid passing implementations.
In the training setup, for each coding task, the rollout policy generates multiple trajectories. The final patch from each trajectory is evaluated using executable tests to yield a binary outcome reward. Let Gx={τi}i=1n denote the group of valid trajectories. The authors use the mean-centered outcome advantage Ai=Ri−Rˉ, where Rˉ=n−1∑j=1nRj. To ensure both successful and failed trajectories are present, dynamic sampling is adopted so that 0<Rˉ<1. All passing trajectories initially receive the same positive advantage 1−Rˉ, which fails to differentiate implementation quality.
To address this, the authors introduce groupwise agentic grading. The grader jointly examines all rollouts in a shared workspace containing the task specification, repository, complete trajectories, submitted patches, and test outputs. This comparison helps identify the strongest solutions and exposes ineffective strategies. The agentic grader gathers evidence iteratively, reviewing turn-by-turn summaries, reading relevant trajectory portions, and cross-checking patches against repository code and test logs. It can also run targeted checks. Quality is assessed across five criteria: approach suitability, precision, minimality, side effects, and codebase consistency. Passing solutions are checked for leaked or external answers, and confirmed hacks receive zero reward. The remaining candidates are scored and mapped to three quality tiers, which determine a discount factor fi∈(0,1] on their positive advantage.
Following grading, sum-preserving advantage redistribution is applied. Simply applying the discount factors would yield A~i=fiAi for passing trajectories, removing positive credit without a corresponding change on the negative side. This deficit weakens the reinforcement of successful trajectories. To preserve the total positive advantage S+=∑i∈PAi assigned to passing trajectories, the authors apply a common rescaling factor:
λ=∑j∈PfjAjS+The redistributed advantage is then:
Ai⋆={λfiAi,Ai,i∈P,i∈F.This ensures the total positive advantage is restored and the change is zero-sum over the passing subset, preserving both the ordering and relative strength of quality preferences.
During online training integration, grading runs asynchronously with rollout collection. The system checks the validity of grading results, falling back to original outcome advantages if necessary. Confirmed reliance on external solutions resets the affected reward to zero before group statistics are recomputed. The resulting sequence-level advantages are broadcast to the model-generated response tokens to supervise the RL policy update.
Experiment
These experiments evaluate Gagar on code-only reinforcement learning with MiMo-V2.6-Flash and on industrial-scale mixed-task RL with Flash and Pro, using DeepSWE v1.1 and SWE-bench Pro as benchmarks. The main comparisons show that Gagar improves long-horizon code agent performance, keeps trajectories shorter and more efficient, and yields higher implementation quality in external model review. An ablation confirms that sum-preserving redistribution is important for training stability, since downweighting passing trajectories without restoring credit leads to unstable dynamics and degraded downstream behavior. Industrial-scale runs further demonstrate competitive code agent performance relative to larger frontier models.
After industrial-scale mixed-task RL with Gagar, MiMo-V2.6-Flash and MiMo-V2.6-Pro achieve coding benchmark results competitive with strong external baselines. MiMo-V2.6-Pro outperforms Kimi K3 on DeepSWE despite having substantially fewer total parameters and surpasses GPT-5.6 Sol on SWE-bench Pro. Its DeepSWE score approaches GPT-5.6 Sol and Claude Opus 5, while it trails Claude Opus 5 on SWE-bench Pro. MiMo-V2.6-Pro outperforms Kimi K3 on DeepSWE despite using substantially fewer total parameters. MiMo-V2.6-Pro surpasses GPT-5.6 Sol on SWE-bench Pro. MiMo-V2.6-Pro approaches GPT-5.6 Sol and Claude Opus 5 on DeepSWE but remains below Claude Opus 5 on SWE-bench Pro. MiMo-V2.6-Flash achieves DeepSWE and SWE-bench Pro scores close to larger external models.
This experiment evaluates MiMo-V2.6-Flash and MiMo-V2.6-Pro after industrial-scale mixed-task RL with Gagar on coding benchmarks. The results show that MiMo-V2.6-Pro is competitive with strong external models, outperforming Kimi K3 on DeepSWE despite using substantially fewer parameters and surpassing GPT-5.6 Sol on SWE-bench Pro, while approaching GPT-5.6 Sol and Claude Opus 5 on DeepSWE but remaining below Claude Opus 5 on SWE-bench Pro. MiMo-V2.6-Flash also achieves DeepSWE and SWE-bench Pro scores close to larger external models.