HyperAIHyperAI

Command Palette

Search for a command to run...

On-Policyoder Off-Policy-Lernen? Eine systematische Studie zur Destillationsdynamik

Julianna Piskorz Antonin Berthon Mihaela van der Schaar

Zusammenfassung

Es wird argumentiert, dass On-Policy-Lernen katastrophales Vergessen reduziert, sparsamere Parameteraktualisierungen erzeugt und die Generalisierung verbessert. Bestehende Vergleiche zwischen überwachtem Fine-Tuning und Reinforcement Learning variieren jedoch viele Faktoren gleichzeitig, wodurch sich der Beitrag der Rollout-Policy nur schwer isolieren lässt. Wir untersuchen den Effekt der Rollout-Policy in einem kontrollierten Strong-to-Weak-Destillationssetting, indem wir Rollout-Policy, Token-Level-KL-Richtung und Lernrate unabhängig voneinander über die Modellfamilien Llama 3 und Qwen2.5 sowie über Reasoning-Aufgaben aus wissenschaftlichen, medizinischen und arithmetischen Domänen hinweg variieren. Unsere Analyse zeigt ein nuanciertes Bild der Destillationsdynamik, in dem die Rollout-Policy nicht notwendigerweise eine zentrale Rolle spielt. Stattdessen prägt die Token-Level-KL-Richtung deutlicher die Aufgabenleistung und die Output-Abdeckung, während die Lernrate das Vergessen und die Update-Sparsamkeit steuert. Die Analyse der KL-Gradienten und Experimente entlang eines kontinuierlichen Student-Teacher-Rollout-Policy-Spektrums erklären dieses Muster: Forward-KL ist bemerkenswert robust gegenüber der Rollout-Policy und bleibt trotz Änderungen der Rollout-Policy stabil und leistungsstark, während Reverse-KL deutlich empfindlicher ist und studentgenerierte Rollouts bevorzugt. On-Policy-Daten verbessern dennoch die Generalisierung auf schwierigere Varianten der arithmetischen Countdown-Aufgabe unter beiden KL-Richtungen, obwohl dieser Vorteil nach anschließendem RLVR nicht zuverlässig bestehen bleibt. Unsere allgemeineren Schlussfolgerungen bleiben robust, wenn Gradient-Clipping entfernt wird, gesampelte KL-Schätzer verwendet werden und auf Aufgaben trainiert wird, die längere Reasoning-Ketten erfordern. Insgesamt stellen unsere Ergebnisse die Auffassung infrage, dass On-Policy-Rollouts inhärent vorzuziehen sind, und zeigen, dass ihr Wert entscheidend vom Optimierungsziel, dem Evaluierungssetting und den Optimierungs-Hyperparametern abhängt.

One-sentence Summary

Researchers at the University of Cambridge systematically investigate strong-to-weak distillation dynamics across Llama 3 and Qwen2.5 model families on reasoning tasks spanning scientific, medical, and arithmetic domains, independently varying rollout policy, token-level \mathrm{\mathrm{KL\mathrm{KL}KL}} direction, and learning rate, and find that forward KL remains robust to rollout policy while reverse KL favors student-generated rollouts, and that on-policy data improves generalization to harder Countdown arithmetic variants although this advantage does not reliably persist after subsequent RLVR.

Key Contributions

  • This paper presents a controlled strong-to-weak distillation study that independently varies rollout policy, token-level KL direction, and learning rate across Llama 3 and Qwen2.5 on scientific, medical, and arithmetic reasoning tasks. The results indicate that token-level KL direction more clearly shapes task performance and output coverage, while learning rate governs forgetting and update sparsity, and rollout policy does not necessarily play a central role.
  • Through KL gradient analysis and experiments along a continuous student-teacher rollout-policy spectrum, the paper shows that forward KL is robust to rollout policy changes, whereas reverse KL is more sensitive and favors student-generated rollouts.
  • On-policy data improves generalization to harder Countdown arithmetic variants under both KL directions, but this advantage does not reliably persist after subsequent RLVR. These conclusions remain robust without gradient clipping, with sampled KL estimators, and on tasks requiring longer reasoning chains, challenging the view that on-policy rollouts are inherently preferable.

Introduction

Post-training is central to developing reasoning capabilities in large language models, and strong-to-weak distillation has become a common way to transfer reasoning behavior from larger teachers to smaller students. On-policy rollouts have been credited with reducing catastrophic forgetting, producing sparser parameter updates, and improving generalization, but prior comparisons often confound rollout policy with changes in the training objective, supervision density, optimization procedure, and learning rate. The authors address this by using strong-to-weak distillation as a controlled testbed, independently varying rollout policy, KL direction, and learning rate across Llama 3 and Qwen2.5 models on scientific, medical, and arithmetic reasoning tasks. They find no consistent advantage from on-policy distillation in final in-distribution accuracy, catastrophic forgetting, or parameter-update sparsity; KL direction more strongly determines task accuracy and coverage, while learning rate governs forgetting and sparsity. Forward KL is robust to rollout policy, whereas reverse KL is substantially more sensitive and favors on-policy rollouts.

Method

5.1 Logit Gradients Reveal Differing Sensitivity to the Rollout Policy

To understand how the rollout policy affects optimization, the authors analyze token-level gradients of forward and reverse KL divergences. Let (zSθ)v,v∈V(z_S^\theta)_v, v \in \mathcal{V}(zSθ​)v​,v∈V denote the logit values produced by the student model at a fixed prefix, with πSθ(v)=softmax(zSθ)v\pi_S^\theta(v) = \mathrm{softmax}(z_S^\theta)_vπSθ​(v)=softmax(zSθ​)v​. Then the parameter gradients can be described as follows:

∇θDF−KL=∑v∈V(πSθ(v)−πT(v))∇θ(zSθ)v,\nabla_{\theta} D_{\mathrm{F-KL}} = \sum_{v \in \mathcal{V}} (\pi_S^\theta(v) - \pi_T(v)) \nabla_{\theta} (z_S^\theta)_v,∇θ​DF−KL​=v∈V∑​(πSθ​(v)−πT​(v))∇θ​(zSθ​)v​, ∇θDR−KL=∑v∈VπSθ(v)[log⁡πSθ(v)πT(v)−DR−KL]∇θ(zSθ)v.\nabla_{\theta} D_{\mathrm{R-KL}} = \sum_{v \in \mathcal{V}} \pi_S^\theta(v) \left[ \log \frac{\pi_S^\theta(v)}{\pi_T(v)} - D_{\mathrm{R-KL}} \right] \nabla_{\theta} (z_S^\theta)_v.∇θ​DR−KL​=v∈V∑​πSθ​(v)[logπT​(v)πSθ​(v)​−DR−KL​]∇θ​(zSθ​)v​.

These expressions reveal an important asymmetry. The forward-KL derivative with respect to the student logits is πSθ(v)−πT(v)\pi_S^\theta(v) - \pi_T(v)πSθ​(v)−πT​(v); it is therefore nonzero whenever the teacher and student next-token distributions differ, and each coordinate lies in [−1,1][-1, 1][−1,1]. Consequently, under bounded student-logit Jacobians ∇θ(zSθ)\nabla_\theta (z_S^\theta)∇θ​(zSθ​), the difference between forward-KL updates induced by two rollout policies is bounded linearly by the total-variation distance between the prefix distributions they induce. Hence, small changes in the trajectories generated by the rollout policy produce proportionally small changes in the forward-KL gradient.

For reverse KL, however, even a small rollout change can produce an arbitrarily large gradient change when the teacher and student assign very different probabilities to some tokens. Its logit derivative is weighted by the student probability πSθ(v)\pi_S^\theta(v)πSθ​(v), so it vanishes as the student probability approaches zero, even when the teacher assigns substantial probability to that token. Reverse KL may therefore struggle to recover teacher modes omitted by the student, reflecting its mode-seeking behavior. Conversely, when the student assigns appreciable probability to a token that the teacher considers extremely unlikely, the log ratio log⁡πSθ(v)/πT(v)\log \pi_S^\theta(v) / \pi_T(v)logπSθ​(v)/πT​(v) can become arbitrarily large, potentially producing sharp, high-variance updates that can destabilize training. Unlike forward KL, reverse KL admits no bound dependent solely on the rollout distance and the student-logit Jacobian.

These properties suggest that reverse KL is more sensitive to the rollout policy. While forward KL provides signal at any visited prefix where the policies disagree, reverse KL emphasizes student-supported, teacher-disfavored tokens and therefore depends more strongly on visiting the student's own prefix distribution.

Experiment

The study runs controlled distillation experiments comparing on-policy and off-policy rollouts while varying token-level KL direction and learning rate across medical, scientific, arithmetic, and longer math reasoning tasks, evaluating task accuracy, catastrophic forgetting, and parameter-update sparsity. The results show that rollout policy provides no consistent advantage for in-distribution performance, forgetting, or sparsity, whereas forward KL is more robust than reverse KL, and learning rate is the main factor governing forgetting and sparsity. A rollout-policy spectrum confirms that forward KL is relatively insensitive to the rollout distribution while reverse KL benefits from student-favored rollouts, and further analyses show that on-policy data can improve generalization to harder tasks and reduce incidental teacher-style transfer, though this advantage does not reliably persist after reinforcement learning. Ablations with sampled KL, no gradient clipping, and longer rollouts indicate that the main conclusions generalize, with only a suggestive on-policy benefit in the long-rollout reverse-KL setting.


KI mit KI entwickeln

Von der Idee bis zum Launch – beschleunigen Sie Ihre KI-Entwicklung mit kostenlosem KI-Co-Coding, sofort einsatzbereiter Umgebung und bestem GPU-Preis.

KI-gestütztes kollaboratives Programmieren
Sofort einsatzbereite GPUs
Die besten Preise

HyperAI Newsletters

Abonnieren Sie unsere neuesten Updates
Wir werden die neuesten Updates der Woche in Ihren Posteingang liefern um neun Uhr jeden Montagmorgen
Unterstützt von MailChimp