HyperAIHyperAI

Command Palette

Search for a command to run...

المعلّم اتجاهٌ لا وجهة: استقراء بواقي التمثيلات المستحثة بالتعلم المعزز في التقطير وفق السياسة

Hao Li MeiJia Chen Weijie Ren Donghan Li Zijun Tian Jingchun Huang Naibo Wang

الملخص

التقطير وفق السياسة (OPD) يدرّب طالبًا على مطابقة توزيعات الرمز التالي للمعلّم على مسارات الطالب نفسه، وقد حقق مكاسب تجريبية كبيرة. وتتيح الصيغ المعمَّمة للطالب تجاوزَ المعلّم من خلال استقراء مكافأة ضمنية في فضاء المخرجات. غير أنّ رأس نموذج اللغة يوهِن هذا التغيير توهينًا متباين الخواص: فجزء كبير من التغيير المرمَّز في الحالات الخفية للمعلّم يصل إلى اللوغيتات بجزء يسير من وزنه، كما أن نسب اللوغاريتمات الاحتمالية للرموز المسحوبة التي يعتمد عليها استقراء فضاء المخرجات تُدخل ضجيجًا يضخّمه الاستقراء، مما يجعل التدريب غير مستقر. نلاحظ أنّ التعلم المعزز (RL) يزيح التمثيلات الداخلية للنموذج نسبةً إلى نقطة التحقق الأساسية، وأنّ اتجاه هذه الإزاحة يمكن قياسه في كل طبقة. وانطلاقًا من هذه الملاحظة، نقترح RIDE (استقراء الاتجاه المستحث بالتعلم المعزز)، الذي يستقرئ التغيير المستحث بالتعلم المعزز مباشرةً في فضاء التمثيل: في كل طبقة وموضع رمز، يحسب RIDE الباقي بين المعلّم ونقطة تحققه السابقة للتعلم المعزز، ويرجع الحالات الخفية للطالب نحو أهداف مُزاحة إلى ما بعد المعلّم على امتداد هذا الباقي. وبشرط وجود مسار مسحوب، يكافئ هذا الانحدارُ تعظيمَ مكافأة اتجاهية خطية يعرّفها الباقي تحت عقوبة تربيعية متمركزة عند المعلّم، وهو ما يوضح صراحةً كيف يحرّك الهدفُ الطالبَ على طول الاتجاه المستحث بالتعلم المعزز مع الحدّ من انحرافه عن المعلّم. عبر أربعة أزواج من معلّمين أساسيين/معلّمين مدرَّبين بالتعلم المعزز تشمل مقاييس ومعماريات وسلالات تدريب مسبق مختلفة، يقارب RIDE المعلّم المدرَّب بالتعلم المعزز أو يتجاوزه في كل زوج، وهو الأسلوب الوحيد الذي يحقق متوسطه ذلك، كما يتفوق باستمرار على استقراء فضاء المخرجات الذي يدهور الطالب كلما كان المعلّم قريبًا من نقطته الأساسية.

One-sentence Summary

The University of Science and Technology of China, Rutgers University, and Zhejiang University propose RIDE (RL-Induced Direction Extrapolation), an on-policy distillation method that extrapolates RL-induced representation residuals between a teacher and its pre-RL checkpoint at every layer and token position, regressing student hidden states beyond the teacher along a directional reward; across four base/RL-teacher pairs spanning different scales, architectures, and pre-training lineages, RIDE approaches or exceeds the RL-trained teachers and consistently outperforms output-space extrapolation, which degrades the student whenever the teacher is close to its base.

Key Contributions

  • The paper reports the observation that reinforcement learning shifts a model’s internal representations relative to its pre-RL checkpoint, and that this shift direction can be measured at every layer.
  • The paper introduces RIDE, a representation-space extrapolation method that computes the residual between the teacher and its pre-RL checkpoint at each layer and token position, then regresses the student’s hidden states toward targets displaced beyond the teacher along that residual; conditioned on a sampled trajectory, this regression is equivalent to maximizing a linear directional reward under a quadratic penalty centered at the teacher.
  • Across four base/RL-teacher pairs spanning different scales, architectures, and pre-training lineages, RIDE approaches or exceeds the RL-trained teacher on every pair and is the only method whose mean does so, and it consistently outperforms output-space extrapolation, which degrades the student whenever the teacher is close to its base.

Introduction

On-policy distillation has become a standard post-training step for transferring reinforcement-learning teacher gains to a student that shares the teacher's pre-RL initialization. Although standard on-policy distillation treats the teacher as the endpoint, prior work shows that students can outperform teachers by extrapolating the teacher-to-reference log-probability ratio or directions in parameter and logit space. The authors argue that output-space extrapolation is limited because the language-model head anisotropically attenuates representation changes, weakly supervising much of the hidden-state residual and leaving earlier layers unconstrained, while sampled-token log-ratios are length-biased and prone to overoptimization. To address this, they propose RIDE, which computes the layerwise hidden-state difference between the frozen RL teacher and its pre-RL checkpoint on student-generated contexts, extrapolates targets beyond the teacher along this RL-induced residual using a single coefficient, and regresses the student's hidden states toward those targets. This provides a deterministic per-rollout representation-space gradient and, across four base/teacher pairs, approaches or exceeds the RL teacher while outperforming on-policy representation distillation and output-space extrapolation on mathematical reasoning benchmarks.

Method

The authors introduce RIDE, a distillation framework built on the premise that an RL-trained teacher specifies both a destination and a direction. The displacement of the teacher from the base policy captures the specific contribution of reinforcement learning, allowing a student initialized at the base policy to continue along this trajectory.

As shown in the framework diagram below:

The method formulates this process through residual extrapolation. Rather than simply matching the teacher, the target is displaced along the RL-induced change. For a given rollout and prefix, the target is defined as:

τt(λ)=sT,t+(λ−1)(sT,t−sB,t)=λsT,t+(1−λ)sB,t,λ≥0.\tau_t(\lambda) = s_{T,t} + (\lambda - 1)(s_{T,t} - s_{B,t}) = \lambda s_{T,t} + (1 - \lambda)s_{B,t}, \quad \lambda \ge 0.τt​(λ)=sT,t​+(λ−1)(sT,t​−sB,t​)=λsT,t​+(1−λ)sB,t​,λ≥0.

When λ=1\lambda = 1λ=1, the target reverts to the teacher. For λ>1\lambda > 1λ>1, the target lies beyond the teacher, prompting the student to extrapolate the change initiated by RL. Because the difference between the teacher and base representations is recomputed for every rollout and prefix, the direction dynamically reflects the RL-induced change on the specific contexts the student visits.

The authors instantiate this extrapolation in the hidden-state space rather than the output log-probability space. For a student rollout, the frozen base and teacher models are evaluated on the same prefix. The RL-induced residual at layer lll and position ttt is computed as:

Δt(l)=hT,t(l)−hB,t(l).\Delta_t^{(l)} = h_{T,t}^{(l)} - h_{B,t}^{(l)}.Δt(l)​=hT,t(l)​−hB,t(l)​.

Since both models process the identical tokens, this residual isolates what the RL training altered in the processing of that specific context. Applying the extrapolation formula to these hidden states yields the RIDE target:

ht⋆(l)=hT,t(l)+(λ−1)Δt(l)=λhT,t(l)+(1−λ)hB,t(l).h_t^{\star(l)} = h_{T,t}^{(l)} + (\lambda - 1)\Delta_t^{(l)} = \lambda h_{T,t}^{(l)} + (1 - \lambda)h_{B,t}^{(l)}.ht⋆(l)​=hT,t(l)​+(λ−1)Δt(l)​=λhT,t(l)​+(1−λ)hB,t(l)​.

The training objective is a representation-space regression toward this extrapolated target:

Lλ(θ)=Ex,y^∼πθ[1∣Llayer∣∑l∈Llayer1M∑t=1Tmt1d∥hθ,t(l)−sg(ht⋆(l))∥22],\mathcal{L}_\lambda(\theta) = \mathbb{E}_{x, \hat{y} \sim \pi_\theta} \Big[ \frac{1}{|\mathcal{L}_{\text{layer}}|} \sum_{l \in \mathcal{L}_{\text{layer}}} \frac{1}{M} \sum_{t=1}^T m_t \frac{1}{d} \| h_{\theta,t}^{(l)} - \text{sg}(h_t^{\star(l)}) \|_2^2 \Big],Lλ​(θ)=Ex,y^​∼πθ​​[∣Llayer​∣1​l∈Llayer​∑​M1​t=1∑T​mt​d1​∥hθ,t(l)​−sg(ht⋆(l)​)∥22​],

where M=∑tmtM = \sum_t m_tM=∑t​mt​. The authors supervise all layers and the final response positions where student-teacher disagreement concentrates.

Measuring the displacement in hidden states is critical for two reasons. First, the output head attenuates the residual anisotropically. A logit target constrains the hidden state weakly along the directions that the head amplifies least, which is precisely where the RL-induced residual concentrates. Second, extrapolation in the output space amplifies sampling noise, with the variance scaling quadratically with the extrapolation factor. In contrast, the hidden-state residual is computed deterministically from two forward passes, ensuring the conditional variance of the gradient remains zero regardless of the extrapolation magnitude.

From an optimization perspective, the regression objective can be interpreted as reward maximization under a proximity constraint. Minimizing the loss with respect to the student hidden state is equivalent to maximizing:

(λ−1)r(h)−12∥h−hT∥22,(\lambda - 1) r(h) - \frac{1}{2} \| h - h_T \|_2^2,(λ−1)r(h)−21​∥h−hT​∥22​,

where r(h)=⟨h−hT,Δ⟩r(h) = \langle h - h_T, \Delta \rangler(h)=⟨h−hT​,Δ⟩. Here, the reward is the inner product of the student displacement from the teacher with the RL-induced residual, and the penalty is quadratic in the same displacement. The parameter λ−1\lambda - 1λ−1 weights the reward against the penalty, ensuring that at λ=1\lambda = 1λ=1, the reward weight vanishes and the method recovers standard representation matching.

Experiment

Across four base/RL-teacher pairs sharing architecture and trained on DAPO-Math-17K prompts, RIDE was evaluated against output-space and representation-space baselines using Avg@16 on AIME 2024, AIME 2025, and AIMO, where it was the only method to exceed its teacher on every pair. The coefficient sweep showed RIDE improves around λ=1.25 and degrades gracefully, while output-space extrapolation is harmed by λ>1 due to sampled-advantage variance. Mechanistic and ablation experiments validate that the RL-induced residual direction is the active ingredient: the representation-space gradient is deterministic, the signal survives head attenuation, and replacing the residual with random, reversed, mismatched-origin, or trajectory-mismatched directions removes the improvement.

Across four base/RL-teacher pairs, RIDE is the only evaluated distillation method whose mean score exceeds the teacher on every pair, while teacher-matching baselines and output-space extrapolation remain below the teacher. ExOPD consistently underperforms, with the largest damage on pairs where the teacher-to-base gap is small. The margin between RIDE and OPRD isolates the benefit of residual-direction extrapolation under a shared training protocol. RIDE is the only method whose mean score lies above the teacher line on every pair. ExOPD falls below the teacher on every pair and can drop below output-space baselines or the untouched student when the teacher-to-base gap is small. Teacher-matching methods remain below their teachers, and RIDE's advantage over OPRD isolates the effect of targeting the RL-induced residual.

Across four base and RL-teacher pairs, RIDE is the only distillation approach whose mean score surpasses the teacher on every pair, while teacher-matching and output-space extrapolation baselines remain below the teacher. ExOPD consistently underperforms and is most harmful when the teacher-to-base gap is small, sometimes falling below output-space baselines or the untouched student. The comparison between RIDE and OPRD, which share a training protocol, isolates the benefit of targeting the RL-induced residual direction.


بناء الذكاء الاصطناعي بالذكاء الاصطناعي

من الفكرة إلى الإطلاق — سرّع تطوير الذكاء الاصطناعي الخاص بك مع المساعدة البرمجية المجانية بالذكاء الاصطناعي، وبيئة جاهزة للاستخدام، وأفضل أسعار لوحدات معالجة الرسومات.

البرمجة التعاونية باستخدام الذكاء الاصطناعي
وحدات GPU جاهزة للعمل
أفضل الأسعار

HyperAI Newsletters

اشترك في آخر تحديثاتنا
سنرسل لك أحدث التحديثات الأسبوعية إلى بريدك الإلكتروني في الساعة التاسعة من صباح كل يوم اثنين
مدعوم بواسطة MailChimp