HyperAIHyperAI

Command Palette

Search for a command to run...

教師は目的地ではなく方向である:オンポリシー蒸留におけるRL誘導表現残差の外挿

Hao Li MeiJia Chen Weijie Ren Donghan Li Zijun Tian Jingchun Huang Naibo Wang

概要

オンポリシー蒸留(OPD)は、生徒モデル自身の軌跡上で生徒モデルを教師の次トークン分布に一致させるよう学習させ、これまでに大きな経験的利得をもたらしてきた。一般化された変種では、出力空間において暗黙の報酬を外挿することで生徒モデルが教師を上回ることが可能になる。しかし、言語モデルヘッドはこの変化を異方的に減衰させる。教師の隠れ状態に符号化された変化の多くは、その大きさのごく一部しかロジットに到達せず、出力空間での外挿が依拠するサンプルトークンの対数確率比は外挿によって増幅されるノイズを注入し、学習を不安定にする。我々は、強化学習(RL)がモデルの内部表現をそのベースチェックポイントに対して移動させること、およびその移動方向が各層で測定可能であることを観察する。この観察に動機づけられ、我々はRIDE(RL-Induced Direction Extrapolation)を提案する。RIDEはRL誘導の変化を表現空間で直接外挿する。すなわち、すべての層とトークン位置において、RIDEは教師とそのRL前チェックポイントの残差を計算し、生徒モデルの隠れ状態を、この残差に沿って教師を超えた位置にあるターゲットへ回帰させる。サンプルされた軌跡に条件付けた場合、この回帰は、教師を中心とする二次ペナルティの下で残差によって定義される線形方向報酬を最大化することと等価である。これにより、目的関数が生徒モデルをRL誘導方向に沿って移動させつつ、教師からの逸脱を制限する仕組みが明示される。異なる規模、アーキテクチャ、事前学習系統にわたる4組のベース/RL教師ペアにおいて、RIDEはすべてのペアでRL学習済み教師に迫るかこれを上回り、平均でそうなる唯一の手法である。また、出力空間での外挿を一貫して上回る。出力空間外挿は、教師がそのベースに近い場合は常に生徒モデルを劣化させる。

One-sentence Summary

The University of Science and Technology of China, Rutgers University, and Zhejiang University propose RIDE (RL-Induced Direction Extrapolation), an on-policy distillation method that extrapolates RL-induced representation residuals between a teacher and its pre-RL checkpoint at every layer and token position, regressing student hidden states beyond the teacher along a directional reward; across four base/RL-teacher pairs spanning different scales, architectures, and pre-training lineages, RIDE approaches or exceeds the RL-trained teachers and consistently outperforms output-space extrapolation, which degrades the student whenever the teacher is close to its base.

Key Contributions

  • The paper reports the observation that reinforcement learning shifts a model’s internal representations relative to its pre-RL checkpoint, and that this shift direction can be measured at every layer.
  • The paper introduces RIDE, a representation-space extrapolation method that computes the residual between the teacher and its pre-RL checkpoint at each layer and token position, then regresses the student’s hidden states toward targets displaced beyond the teacher along that residual; conditioned on a sampled trajectory, this regression is equivalent to maximizing a linear directional reward under a quadratic penalty centered at the teacher.
  • Across four base/RL-teacher pairs spanning different scales, architectures, and pre-training lineages, RIDE approaches or exceeds the RL-trained teacher on every pair and is the only method whose mean does so, and it consistently outperforms output-space extrapolation, which degrades the student whenever the teacher is close to its base.

Introduction

On-policy distillation has become a standard post-training step for transferring reinforcement-learning teacher gains to a student that shares the teacher's pre-RL initialization. Although standard on-policy distillation treats the teacher as the endpoint, prior work shows that students can outperform teachers by extrapolating the teacher-to-reference log-probability ratio or directions in parameter and logit space. The authors argue that output-space extrapolation is limited because the language-model head anisotropically attenuates representation changes, weakly supervising much of the hidden-state residual and leaving earlier layers unconstrained, while sampled-token log-ratios are length-biased and prone to overoptimization. To address this, they propose RIDE, which computes the layerwise hidden-state difference between the frozen RL teacher and its pre-RL checkpoint on student-generated contexts, extrapolates targets beyond the teacher along this RL-induced residual using a single coefficient, and regresses the student's hidden states toward those targets. This provides a deterministic per-rollout representation-space gradient and, across four base/teacher pairs, approaches or exceeds the RL teacher while outperforming on-policy representation distillation and output-space extrapolation on mathematical reasoning benchmarks.

Method

The authors introduce RIDE, a distillation framework built on the premise that an RL-trained teacher specifies both a destination and a direction. The displacement of the teacher from the base policy captures the specific contribution of reinforcement learning, allowing a student initialized at the base policy to continue along this trajectory.

As shown in the framework diagram below:

The method formulates this process through residual extrapolation. Rather than simply matching the teacher, the target is displaced along the RL-induced change. For a given rollout and prefix, the target is defined as:

τt(λ)=sT,t+(λ−1)(sT,t−sB,t)=λsT,t+(1−λ)sB,t,λ≥0.\tau_t(\lambda) = s_{T,t} + (\lambda - 1)(s_{T,t} - s_{B,t}) = \lambda s_{T,t} + (1 - \lambda)s_{B,t}, \quad \lambda \ge 0.τt​(λ)=sT,t​+(λ−1)(sT,t​−sB,t​)=λsT,t​+(1−λ)sB,t​,λ≥0.

When λ=1\lambda = 1λ=1, the target reverts to the teacher. For λ>1\lambda > 1λ>1, the target lies beyond the teacher, prompting the student to extrapolate the change initiated by RL. Because the difference between the teacher and base representations is recomputed for every rollout and prefix, the direction dynamically reflects the RL-induced change on the specific contexts the student visits.

The authors instantiate this extrapolation in the hidden-state space rather than the output log-probability space. For a student rollout, the frozen base and teacher models are evaluated on the same prefix. The RL-induced residual at layer lll and position ttt is computed as:

Δt(l)=hT,t(l)−hB,t(l).\Delta_t^{(l)} = h_{T,t}^{(l)} - h_{B,t}^{(l)}.Δt(l)​=hT,t(l)​−hB,t(l)​.

Since both models process the identical tokens, this residual isolates what the RL training altered in the processing of that specific context. Applying the extrapolation formula to these hidden states yields the RIDE target:

ht⋆(l)=hT,t(l)+(λ−1)Δt(l)=λhT,t(l)+(1−λ)hB,t(l).h_t^{\star(l)} = h_{T,t}^{(l)} + (\lambda - 1)\Delta_t^{(l)} = \lambda h_{T,t}^{(l)} + (1 - \lambda)h_{B,t}^{(l)}.ht⋆(l)​=hT,t(l)​+(λ−1)Δt(l)​=λhT,t(l)​+(1−λ)hB,t(l)​.

The training objective is a representation-space regression toward this extrapolated target:

Lλ(θ)=Ex,y^∼πθ[1∣Llayer∣∑l∈Llayer1M∑t=1Tmt1d∥hθ,t(l)−sg(ht⋆(l))∥22],\mathcal{L}_\lambda(\theta) = \mathbb{E}_{x, \hat{y} \sim \pi_\theta} \Big[ \frac{1}{|\mathcal{L}_{\text{layer}}|} \sum_{l \in \mathcal{L}_{\text{layer}}} \frac{1}{M} \sum_{t=1}^T m_t \frac{1}{d} \| h_{\theta,t}^{(l)} - \text{sg}(h_t^{\star(l)}) \|_2^2 \Big],Lλ​(θ)=Ex,y^​∼πθ​​[∣Llayer​∣1​l∈Llayer​∑​M1​t=1∑T​mt​d1​∥hθ,t(l)​−sg(ht⋆(l)​)∥22​],

where M=∑tmtM = \sum_t m_tM=∑t​mt​. The authors supervise all layers and the final response positions where student-teacher disagreement concentrates.

Measuring the displacement in hidden states is critical for two reasons. First, the output head attenuates the residual anisotropically. A logit target constrains the hidden state weakly along the directions that the head amplifies least, which is precisely where the RL-induced residual concentrates. Second, extrapolation in the output space amplifies sampling noise, with the variance scaling quadratically with the extrapolation factor. In contrast, the hidden-state residual is computed deterministically from two forward passes, ensuring the conditional variance of the gradient remains zero regardless of the extrapolation magnitude.

From an optimization perspective, the regression objective can be interpreted as reward maximization under a proximity constraint. Minimizing the loss with respect to the student hidden state is equivalent to maximizing:

(λ−1)r(h)−12∥h−hT∥22,(\lambda - 1) r(h) - \frac{1}{2} \| h - h_T \|_2^2,(λ−1)r(h)−21​∥h−hT​∥22​,

where r(h)=⟨h−hT,Δ⟩r(h) = \langle h - h_T, \Delta \rangler(h)=⟨h−hT​,Δ⟩. Here, the reward is the inner product of the student displacement from the teacher with the RL-induced residual, and the penalty is quadratic in the same displacement. The parameter λ−1\lambda - 1λ−1 weights the reward against the penalty, ensuring that at λ=1\lambda = 1λ=1, the reward weight vanishes and the method recovers standard representation matching.

Experiment

Across four base/RL-teacher pairs sharing architecture and trained on DAPO-Math-17K prompts, RIDE was evaluated against output-space and representation-space baselines using Avg@16 on AIME 2024, AIME 2025, and AIMO, where it was the only method to exceed its teacher on every pair. The coefficient sweep showed RIDE improves around λ=1.25 and degrades gracefully, while output-space extrapolation is harmed by λ>1 due to sampled-advantage variance. Mechanistic and ablation experiments validate that the RL-induced residual direction is the active ingredient: the representation-space gradient is deterministic, the signal survives head attenuation, and replacing the residual with random, reversed, mismatched-origin, or trajectory-mismatched directions removes the improvement.

Across four base/RL-teacher pairs, RIDE is the only evaluated distillation method whose mean score exceeds the teacher on every pair, while teacher-matching baselines and output-space extrapolation remain below the teacher. ExOPD consistently underperforms, with the largest damage on pairs where the teacher-to-base gap is small. The margin between RIDE and OPRD isolates the benefit of residual-direction extrapolation under a shared training protocol. RIDE is the only method whose mean score lies above the teacher line on every pair. ExOPD falls below the teacher on every pair and can drop below output-space baselines or the untouched student when the teacher-to-base gap is small. Teacher-matching methods remain below their teachers, and RIDE's advantage over OPRD isolates the effect of targeting the RL-induced residual.

Across four base and RL-teacher pairs, RIDE is the only distillation approach whose mean score surpasses the teacher on every pair, while teacher-matching and output-space extrapolation baselines remain below the teacher. ExOPD consistently underperforms and is most harmful when the teacher-to-base gap is small, sometimes falling below output-space baselines or the untouched student. The comparison between RIDE and OPRD, which share a training protocol, isolates the benefit of targeting the RL-induced residual direction.


AIでAIを構築

アイデアからローンチまで — 無料のAIコーディング支援、すぐに使える環境、最高のGPU価格でAI開発を加速。

AI コーディング補助
すぐに使える GPU
最適な料金体系

HyperAI Newsletters

最新情報を購読する
北京時間 毎週月曜日の午前9時 に、その週の最新情報をメールでお届けします
メール配信サービスは MailChimp によって提供されています