Command Palette
Search for a command to run...
Der Lehrer ist eine Richtung, kein Ziel: Extrapolation RL-induzierter Repräsentationsresiduen in der On-Policy-Destillation
Der Lehrer ist eine Richtung, kein Ziel: Extrapolation RL-induzierter Repräsentationsresiduen in der On-Policy-Destillation
Hao Li MeiJia Chen Weijie Ren Donghan Li Zijun Tian Jingchun Huang Naibo Wang
Zusammenfassung
On-Policy-Destillation (OPD) trainiert einen Schüler darauf, die Next-Token-Verteilungen des Lehrers auf den eigenen Trajektorien des Schülers abzugleichen, und hat zu erheblichen empirischen Verbesserungen geführt. Verallgemeinerte Varianten ermöglichen es dem Schüler, den Lehrer zu übertreffen, indem sie eine implizite Belohnung im Ausgaberaum extrapolieren. Der Sprachmodell-Kopf dämpft diese Änderung jedoch anisotrop: Ein großer Teil der in den verborgenen Zuständen des Lehrers kodierten Änderung erreicht die Logits nur mit einem kleinen Bruchteil ihrer Stärke, und die Log-Wahrscheinlichkeitsverhältnisse gesampelter Token, auf die sich die Ausgaberaum-Extrapolation stützt, injizieren Rauschen, das durch die Extrapolation verstärkt wird, wodurch das Training instabil wird. Wir beobachten, dass Reinforcement Learning (RL) die internen Repräsentationen eines Modells relativ zu seinem Basis-Checkpoint verschiebt und dass die Richtung dieser Verschiebung in jeder Schicht gemessen werden kann. Motiviert durch diese Beobachtung schlagen wir RIDE (RL-Induced Direction Extrapolation) vor, das die RL-induzierte Änderung direkt im Repräsentationsraum extrapoliert: In jeder Schicht und an jeder Token-Position berechnet RIDE das Residuum zwischen dem Lehrer und seinem Prä-RL-Checkpoint und führt eine Regression der verborgenen Zustände des Schülers auf Ziele durch, die entlang dieses Residuums über den Lehrer hinaus verschoben sind. Bedingt auf eine gesampelte Trajektorie ist diese Regression äquivalent zur Maximierung einer linearen Richtungsbelohnung, die durch das Residuum definiert wird, unter einer quadratischen Strafe, die um den Lehrer zentriert ist; dies verdeutlicht, wie die Zielfunktion den Schüler entlang der RL-induzierten Richtung bewegt, während sie seine Abweichung vom Lehrer begrenzt. Über vier Basis-/RL-Lehrer-Paare hinweg, die unterschiedliche Skalierungen, Architekturen und Vortrainingslinien abdecken, erreicht RIDE auf jedem Paar den RL-trainierten Lehrer oder übertrifft ihn und ist die einzige Methode, deren Mittelwert dies über alle Paare hinweg erreicht; zudem übertrifft RIDE durchgängig die Ausgaberaum-Extrapolation, die den Schüler immer dann verschlechtert, wenn der Lehrer nahe an seiner Basis liegt.
One-sentence Summary
The University of Science and Technology of China, Rutgers University, and Zhejiang University propose RIDE (RL-Induced Direction Extrapolation), an on-policy distillation method that extrapolates RL-induced representation residuals between a teacher and its pre-RL checkpoint at every layer and token position, regressing student hidden states beyond the teacher along a directional reward; across four base/RL-teacher pairs spanning different scales, architectures, and pre-training lineages, RIDE approaches or exceeds the RL-trained teachers and consistently outperforms output-space extrapolation, which degrades the student whenever the teacher is close to its base.
Key Contributions
- The paper reports the observation that reinforcement learning shifts a model’s internal representations relative to its pre-RL checkpoint, and that this shift direction can be measured at every layer.
- The paper introduces RIDE, a representation-space extrapolation method that computes the residual between the teacher and its pre-RL checkpoint at each layer and token position, then regresses the student’s hidden states toward targets displaced beyond the teacher along that residual; conditioned on a sampled trajectory, this regression is equivalent to maximizing a linear directional reward under a quadratic penalty centered at the teacher.
- Across four base/RL-teacher pairs spanning different scales, architectures, and pre-training lineages, RIDE approaches or exceeds the RL-trained teacher on every pair and is the only method whose mean does so, and it consistently outperforms output-space extrapolation, which degrades the student whenever the teacher is close to its base.
Introduction
On-policy distillation has become a standard post-training step for transferring reinforcement-learning teacher gains to a student that shares the teacher's pre-RL initialization. Although standard on-policy distillation treats the teacher as the endpoint, prior work shows that students can outperform teachers by extrapolating the teacher-to-reference log-probability ratio or directions in parameter and logit space. The authors argue that output-space extrapolation is limited because the language-model head anisotropically attenuates representation changes, weakly supervising much of the hidden-state residual and leaving earlier layers unconstrained, while sampled-token log-ratios are length-biased and prone to overoptimization. To address this, they propose RIDE, which computes the layerwise hidden-state difference between the frozen RL teacher and its pre-RL checkpoint on student-generated contexts, extrapolates targets beyond the teacher along this RL-induced residual using a single coefficient, and regresses the student's hidden states toward those targets. This provides a deterministic per-rollout representation-space gradient and, across four base/teacher pairs, approaches or exceeds the RL teacher while outperforming on-policy representation distillation and output-space extrapolation on mathematical reasoning benchmarks.
Method
The authors introduce RIDE, a distillation framework built on the premise that an RL-trained teacher specifies both a destination and a direction. The displacement of the teacher from the base policy captures the specific contribution of reinforcement learning, allowing a student initialized at the base policy to continue along this trajectory.
As shown in the framework diagram below:
The method formulates this process through residual extrapolation. Rather than simply matching the teacher, the target is displaced along the RL-induced change. For a given rollout and prefix, the target is defined as:
τt(λ)=sT,t+(λ−1)(sT,t−sB,t)=λsT,t+(1−λ)sB,t,λ≥0.When λ=1, the target reverts to the teacher. For λ>1, the target lies beyond the teacher, prompting the student to extrapolate the change initiated by RL. Because the difference between the teacher and base representations is recomputed for every rollout and prefix, the direction dynamically reflects the RL-induced change on the specific contexts the student visits.
The authors instantiate this extrapolation in the hidden-state space rather than the output log-probability space. For a student rollout, the frozen base and teacher models are evaluated on the same prefix. The RL-induced residual at layer l and position t is computed as:
Δt(l)=hT,t(l)−hB,t(l).Since both models process the identical tokens, this residual isolates what the RL training altered in the processing of that specific context. Applying the extrapolation formula to these hidden states yields the RIDE target:
ht⋆(l)=hT,t(l)+(λ−1)Δt(l)=λhT,t(l)+(1−λ)hB,t(l).The training objective is a representation-space regression toward this extrapolated target:
Lλ(θ)=Ex,y^∼πθ[∣Llayer∣1l∈Llayer∑M1t=1∑Tmtd1∥hθ,t(l)−sg(ht⋆(l))∥22],where M=∑tmt. The authors supervise all layers and the final response positions where student-teacher disagreement concentrates.
Measuring the displacement in hidden states is critical for two reasons. First, the output head attenuates the residual anisotropically. A logit target constrains the hidden state weakly along the directions that the head amplifies least, which is precisely where the RL-induced residual concentrates. Second, extrapolation in the output space amplifies sampling noise, with the variance scaling quadratically with the extrapolation factor. In contrast, the hidden-state residual is computed deterministically from two forward passes, ensuring the conditional variance of the gradient remains zero regardless of the extrapolation magnitude.
From an optimization perspective, the regression objective can be interpreted as reward maximization under a proximity constraint. Minimizing the loss with respect to the student hidden state is equivalent to maximizing:
(λ−1)r(h)−21∥h−hT∥22,where r(h)=⟨h−hT,Δ⟩. Here, the reward is the inner product of the student displacement from the teacher with the RL-induced residual, and the penalty is quadratic in the same displacement. The parameter λ−1 weights the reward against the penalty, ensuring that at λ=1, the reward weight vanishes and the method recovers standard representation matching.
Experiment
Across four base/RL-teacher pairs sharing architecture and trained on DAPO-Math-17K prompts, RIDE was evaluated against output-space and representation-space baselines using Avg@16 on AIME 2024, AIME 2025, and AIMO, where it was the only method to exceed its teacher on every pair. The coefficient sweep showed RIDE improves around λ=1.25 and degrades gracefully, while output-space extrapolation is harmed by λ>1 due to sampled-advantage variance. Mechanistic and ablation experiments validate that the RL-induced residual direction is the active ingredient: the representation-space gradient is deterministic, the signal survives head attenuation, and replacing the residual with random, reversed, mismatched-origin, or trajectory-mismatched directions removes the improvement.
Across four base/RL-teacher pairs, RIDE is the only evaluated distillation method whose mean score exceeds the teacher on every pair, while teacher-matching baselines and output-space extrapolation remain below the teacher. ExOPD consistently underperforms, with the largest damage on pairs where the teacher-to-base gap is small. The margin between RIDE and OPRD isolates the benefit of residual-direction extrapolation under a shared training protocol. RIDE is the only method whose mean score lies above the teacher line on every pair. ExOPD falls below the teacher on every pair and can drop below output-space baselines or the untouched student when the teacher-to-base gap is small. Teacher-matching methods remain below their teachers, and RIDE's advantage over OPRD isolates the effect of targeting the RL-induced residual.
Across four base and RL-teacher pairs, RIDE is the only distillation approach whose mean score surpasses the teacher on every pair, while teacher-matching and output-space extrapolation baselines remain below the teacher. ExOPD consistently underperforms and is most harmful when the teacher-to-base gap is small, sometimes falling below output-space baselines or the untouched student. The comparison between RIDE and OPRD, which share a training protocol, isolates the benefit of targeting the RL-induced residual direction.