HyperAIHyperAI

Command Palette

Search for a command to run...

L’enseignant est une direction, pas une destination : extrapolation des résidus de représentation induits par RL dans la distillation on-policy

Hao Li MeiJia Chen Weijie Ren Donghan Li Zijun Tian Jingchun Huang Naibo Wang

Résumé

La distillation on-policy (OPD) entraîne un élève à aligner, sur ses propres trajectoires, les distributions du jeton suivant de l’enseignant et a produit des gains empiriques substantiels. Des variantes généralisées permettent à l’élève de dépasser l’enseignant en extrapolant une récompense implicite dans l’espace de sortie. La tête du modèle de langue atténue toutefois ce changement de manière anisotrope : une grande partie du changement encodé dans les états cachés de l’enseignant n’atteint les logits qu’à une fraction réduite de son poids, et les rapports de log-probabilités des jetons échantillonnés sur lesquels repose l’extrapolation dans l’espace de sortie injectent un bruit que l’extrapolation amplifie, rendant l’entraînement instable. Nous observons que l’apprentissage par renforcement (RL) déplace les représentations internes d’un modèle par rapport à son checkpoint de base, et que la direction de ce déplacement peut être mesurée à chaque couche. Motivés par cette observation, nous proposons RIDE (RL-Induced Direction Extrapolation), qui extrapole le changement induit par RL directement dans l’espace des représentations : à chaque couche et position de jeton, RIDE calcule le résidu entre l’enseignant et son checkpoint pré-RL et régresse les états cachés de l’élève vers des cibles déplacées au-delà de l’enseignant le long de ce résidu. Conditionnée à une trajectoire échantillonnée, cette régression équivaut à maximiser une récompense directionnelle linéaire définie par le résidu sous une pénalité quadratique centrée sur l’enseignant, ce qui rend explicite la façon dont l’objectif déplace l’élève le long de la direction induite par RL tout en limitant son écart par rapport à l’enseignant. Sur quatre paires base/enseignant-RL couvrant différentes échelles, architectures et lignées de pré-entraînement, RIDE approche ou dépasse l’enseignant entraîné par RL sur chaque paire et est la seule méthode dont la moyenne y parvient ; il surpasse en outre systématiquement l’extrapolation dans l’espace de sortie, qui dégrade l’élève chaque fois que l’enseignant est proche de son checkpoint de base.

One-sentence Summary

The University of Science and Technology of China, Rutgers University, and Zhejiang University propose RIDE (RL-Induced Direction Extrapolation), an on-policy distillation method that extrapolates RL-induced representation residuals between a teacher and its pre-RL checkpoint at every layer and token position, regressing student hidden states beyond the teacher along a directional reward; across four base/RL-teacher pairs spanning different scales, architectures, and pre-training lineages, RIDE approaches or exceeds the RL-trained teachers and consistently outperforms output-space extrapolation, which degrades the student whenever the teacher is close to its base.

Key Contributions

  • The paper reports the observation that reinforcement learning shifts a model’s internal representations relative to its pre-RL checkpoint, and that this shift direction can be measured at every layer.
  • The paper introduces RIDE, a representation-space extrapolation method that computes the residual between the teacher and its pre-RL checkpoint at each layer and token position, then regresses the student’s hidden states toward targets displaced beyond the teacher along that residual; conditioned on a sampled trajectory, this regression is equivalent to maximizing a linear directional reward under a quadratic penalty centered at the teacher.
  • Across four base/RL-teacher pairs spanning different scales, architectures, and pre-training lineages, RIDE approaches or exceeds the RL-trained teacher on every pair and is the only method whose mean does so, and it consistently outperforms output-space extrapolation, which degrades the student whenever the teacher is close to its base.

Introduction

On-policy distillation has become a standard post-training step for transferring reinforcement-learning teacher gains to a student that shares the teacher's pre-RL initialization. Although standard on-policy distillation treats the teacher as the endpoint, prior work shows that students can outperform teachers by extrapolating the teacher-to-reference log-probability ratio or directions in parameter and logit space. The authors argue that output-space extrapolation is limited because the language-model head anisotropically attenuates representation changes, weakly supervising much of the hidden-state residual and leaving earlier layers unconstrained, while sampled-token log-ratios are length-biased and prone to overoptimization. To address this, they propose RIDE, which computes the layerwise hidden-state difference between the frozen RL teacher and its pre-RL checkpoint on student-generated contexts, extrapolates targets beyond the teacher along this RL-induced residual using a single coefficient, and regresses the student's hidden states toward those targets. This provides a deterministic per-rollout representation-space gradient and, across four base/teacher pairs, approaches or exceeds the RL teacher while outperforming on-policy representation distillation and output-space extrapolation on mathematical reasoning benchmarks.

Method

The authors introduce RIDE, a distillation framework built on the premise that an RL-trained teacher specifies both a destination and a direction. The displacement of the teacher from the base policy captures the specific contribution of reinforcement learning, allowing a student initialized at the base policy to continue along this trajectory.

As shown in the framework diagram below:

The method formulates this process through residual extrapolation. Rather than simply matching the teacher, the target is displaced along the RL-induced change. For a given rollout and prefix, the target is defined as:

τt(λ)=sT,t+(λ−1)(sT,t−sB,t)=λsT,t+(1−λ)sB,t,λ≥0.\tau_t(\lambda) = s_{T,t} + (\lambda - 1)(s_{T,t} - s_{B,t}) = \lambda s_{T,t} + (1 - \lambda)s_{B,t}, \quad \lambda \ge 0.τt​(λ)=sT,t​+(λ−1)(sT,t​−sB,t​)=λsT,t​+(1−λ)sB,t​,λ≥0.

When λ=1\lambda = 1λ=1, the target reverts to the teacher. For λ>1\lambda > 1λ>1, the target lies beyond the teacher, prompting the student to extrapolate the change initiated by RL. Because the difference between the teacher and base representations is recomputed for every rollout and prefix, the direction dynamically reflects the RL-induced change on the specific contexts the student visits.

The authors instantiate this extrapolation in the hidden-state space rather than the output log-probability space. For a student rollout, the frozen base and teacher models are evaluated on the same prefix. The RL-induced residual at layer lll and position ttt is computed as:

Δt(l)=hT,t(l)−hB,t(l).\Delta_t^{(l)} = h_{T,t}^{(l)} - h_{B,t}^{(l)}.Δt(l)​=hT,t(l)​−hB,t(l)​.

Since both models process the identical tokens, this residual isolates what the RL training altered in the processing of that specific context. Applying the extrapolation formula to these hidden states yields the RIDE target:

ht⋆(l)=hT,t(l)+(λ−1)Δt(l)=λhT,t(l)+(1−λ)hB,t(l).h_t^{\star(l)} = h_{T,t}^{(l)} + (\lambda - 1)\Delta_t^{(l)} = \lambda h_{T,t}^{(l)} + (1 - \lambda)h_{B,t}^{(l)}.ht⋆(l)​=hT,t(l)​+(λ−1)Δt(l)​=λhT,t(l)​+(1−λ)hB,t(l)​.

The training objective is a representation-space regression toward this extrapolated target:

Lλ(θ)=Ex,y^∼πθ[1∣Llayer∣∑l∈Llayer1M∑t=1Tmt1d∥hθ,t(l)−sg(ht⋆(l))∥22],\mathcal{L}_\lambda(\theta) = \mathbb{E}_{x, \hat{y} \sim \pi_\theta} \Big[ \frac{1}{|\mathcal{L}_{\text{layer}}|} \sum_{l \in \mathcal{L}_{\text{layer}}} \frac{1}{M} \sum_{t=1}^T m_t \frac{1}{d} \| h_{\theta,t}^{(l)} - \text{sg}(h_t^{\star(l)}) \|_2^2 \Big],Lλ​(θ)=Ex,y^​∼πθ​​[∣Llayer​∣1​l∈Llayer​∑​M1​t=1∑T​mt​d1​∥hθ,t(l)​−sg(ht⋆(l)​)∥22​],

where M=∑tmtM = \sum_t m_tM=∑t​mt​. The authors supervise all layers and the final response positions where student-teacher disagreement concentrates.

Measuring the displacement in hidden states is critical for two reasons. First, the output head attenuates the residual anisotropically. A logit target constrains the hidden state weakly along the directions that the head amplifies least, which is precisely where the RL-induced residual concentrates. Second, extrapolation in the output space amplifies sampling noise, with the variance scaling quadratically with the extrapolation factor. In contrast, the hidden-state residual is computed deterministically from two forward passes, ensuring the conditional variance of the gradient remains zero regardless of the extrapolation magnitude.

From an optimization perspective, the regression objective can be interpreted as reward maximization under a proximity constraint. Minimizing the loss with respect to the student hidden state is equivalent to maximizing:

(λ−1)r(h)−12∥h−hT∥22,(\lambda - 1) r(h) - \frac{1}{2} \| h - h_T \|_2^2,(λ−1)r(h)−21​∥h−hT​∥22​,

where r(h)=⟨h−hT,Δ⟩r(h) = \langle h - h_T, \Delta \rangler(h)=⟨h−hT​,Δ⟩. Here, the reward is the inner product of the student displacement from the teacher with the RL-induced residual, and the penalty is quadratic in the same displacement. The parameter λ−1\lambda - 1λ−1 weights the reward against the penalty, ensuring that at λ=1\lambda = 1λ=1, the reward weight vanishes and the method recovers standard representation matching.

Experiment

Across four base/RL-teacher pairs sharing architecture and trained on DAPO-Math-17K prompts, RIDE was evaluated against output-space and representation-space baselines using Avg@16 on AIME 2024, AIME 2025, and AIMO, where it was the only method to exceed its teacher on every pair. The coefficient sweep showed RIDE improves around λ=1.25 and degrades gracefully, while output-space extrapolation is harmed by λ>1 due to sampled-advantage variance. Mechanistic and ablation experiments validate that the RL-induced residual direction is the active ingredient: the representation-space gradient is deterministic, the signal survives head attenuation, and replacing the residual with random, reversed, mismatched-origin, or trajectory-mismatched directions removes the improvement.

Across four base/RL-teacher pairs, RIDE is the only evaluated distillation method whose mean score exceeds the teacher on every pair, while teacher-matching baselines and output-space extrapolation remain below the teacher. ExOPD consistently underperforms, with the largest damage on pairs where the teacher-to-base gap is small. The margin between RIDE and OPRD isolates the benefit of residual-direction extrapolation under a shared training protocol. RIDE is the only method whose mean score lies above the teacher line on every pair. ExOPD falls below the teacher on every pair and can drop below output-space baselines or the untouched student when the teacher-to-base gap is small. Teacher-matching methods remain below their teachers, and RIDE's advantage over OPRD isolates the effect of targeting the RL-induced residual.

Across four base and RL-teacher pairs, RIDE is the only distillation approach whose mean score surpasses the teacher on every pair, while teacher-matching and output-space extrapolation baselines remain below the teacher. ExOPD consistently underperforms and is most harmful when the teacher-to-base gap is small, sometimes falling below output-space baselines or the untouched student. The comparison between RIDE and OPRD, which share a training protocol, isolates the benefit of targeting the RL-induced residual direction.


Créer de l'IA avec l'IA

De l'idée au lancement — accélérez votre développement IA avec le co-codage IA gratuit, un environnement prêt à l'emploi et le meilleur prix pour les GPU.

Codage assisté par IA
GPU prêts à l’emploi
Tarifs les plus avantageux

HyperAI Newsletters

Abonnez-vous à nos dernières mises à jour
Nous vous enverrons les dernières mises à jour de la semaine dans votre boîte de réception à neuf heures chaque lundi matin
Propulsé par MailChimp