HyperAIHyperAI

Command Palette

Search for a command to run...

Über die Lehrerzuweisung hinaus: Domänen-normalisierte Multi-Teacher-On-Policy-Destillation

Xin Li Hao Jiang Xin Gao Annan Wang Yuchen Xie Jinghao Guo Xingwei Qu Yichi Zhang Chau Yuen

Zusammenfassung

Bestärkendes Lernen kann ein Sprachmodell in mehrere Spezialisten verwandeln, die jeweils in einer einzelnen Fähigkeit wie Mathematik, Programmieren oder dem Befolgen von Anweisungen exzellent sind; Nutzer benötigen jedoch ein Modell mit all diesen Fähigkeiten. Multi-Teacher-On-Policy-Destillation (MOPD) vereint diese, indem die Spezialisten einen Schüler unterrichten: Der Schüler beantwortet jeden Prompt, und der Spezialist für die Domäne dieses Prompts gibt Feedback zu jedem Token. Dieses Routing entscheidet, welcher Spezialist unterrichtet, aber nicht, wie stark dessen Feedback den gemeinsamen Schüler verändert. Bei Qwen3.5-Modellen in drei Größen stellen wir fest, dass der Schüler von MOPD einen Schüler, der vom besten Einzelspezialisten unterrichtet wurde, nicht übertrifft und nur wenig vom Vorteil des Mathematik-Spezialisten profitiert. Das Feedback ist unausgewogen: Das Feedback zum Befolgen von Anweisungen ist um ein Mehrfaches breiter gestreut als das Mathematik-Feedback und dominiert die Aktualisierungen des Schülers. Wir schlagen Domain-Normalized MOPD (DN-MOPD) vor, das das Routing beibehält und das Feedback jeder Domäne anhand seiner gemessenen Streuung neu skaliert. Auf sechs öffentlichen Benchmarks verbessert DN-MOPD die durchschnittliche Punktzahl gegenüber MOPD bei jeder Größe, über drei zufällige Seeds hinweg und unter zwei Begrenzungen der Antwortlänge, und stellt den größten Teil des verlorenen Mathematik-Gewinns wieder her. Kontrollen mit festen Domänengewichten zeigen, dass der Gewinn hauptsächlich darauf zurückgeht, das Feedback zum Befolgen von Anweisungen herunterzuregeln, statt allein das Mathematik-Feedback hochzuregeln; zudem schneiden feste Gewichte nahe den von DN-MOPD gemessenen Werten vergleichbar ab. Das Kombinieren von Spezialisten erfordert daher nicht nur die Entscheidung, welcher Spezialist unterrichtet, sondern auch, wie stark dessen Feedback gewichtet wird.

One-sentence Summary

Researchers from Nanyang Technological University, Yale University, and University of Manchester propose Domain-Normalized Multi-Teacher On-Policy Distillation (DN-MOPD), which keeps specialist routing but rescales each domain's feedback by its measured spread to correct unbalanced feedback, and on six public benchmarks DN-MOPD outperforms MOPD across model sizes, seeds, and answer-length limits while recovering most of the lost mathematics gains.

Key Contributions

  • In Qwen3.5 models at three sizes, multi-teacher on-policy distillation (MOPD) exhibits unbalanced feedback: instruction-following feedback has several times the spread of mathematics feedback, so the student gains little from the mathematics specialist and does not beat the strongest single-teacher baseline.
  • Domain-Normalized MOPD (DN-MOPD) keeps MOPD's teacher routing and rescales each domain's feedback by the measured spread of its teacher-student log-ratios.
  • On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and two answer-length limits, and recovers most of the lost mathematics gain; fixed-weight controls show the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone.

Introduction

Post-training for large language models increasingly uses domain-specific reinforcement learning, producing separate experts for mathematics, code, and instruction following that must be combined into one deployable model. Multi-teacher on-policy distillation (MOPD) integrates these experts by routing each prompt to its domain expert and using that expert’s token-level log-probabilities on the student’s own response, but routing alone does not control how much each expert’s feedback counts. Fixed or gap-based domain weights ignore differences in feedback spread, so dispersed instruction-following feedback can dominate the shared gradient while mathematics transfer remains weak. The authors introduce Domain-Normalized Multi-Teacher On-Policy Distillation (DN-MOPD), which estimates the spread of teacher–student log-ratios within each domain on every batch and applies a bounded multiplier to bring each domain’s distillation advantages toward the pooled scale. Across Qwen3.5 expert pools at 9B, 4B, and 2B, DN-MOPD improves the six-task average over MOPD and recovers much of the mathematics gain that MOPD loses, without adding extra teacher models, teacher calls, or learned routers.

Method

The authors propose DN-MOPD, a method that retains the domain-label assignment of multi-teacher Online Policy Distillation (OPD) while introducing a domain normalization step. Before the shared student model is updated, the distillation advantages for each domain are rescaled based on their feedback spread relative to the current batch.

The training setup begins with one initial model. Three copies are trained independently via reinforcement learning to become specialists for mathematics, code, and instruction following (IF). A fourth copy serves as the student πu\pi_uπu​, which is trained to acquire all three skills. Each training prompt carries a domain label d∈{math, code, IF}d \in \{\text{math, code, IF}\}d∈{math, code, IF}, and TdT_dTd​ denotes the corresponding domain specialist.

To generate feedback, the student first produces a response y=(y1,…,yn)y = (y_1, \ldots, y_n)y=(y1​,…,yn​) to a prompt. The matching domain teacher then evaluates this response, calculating the token-level distillation advantage by comparing the teacher's token probability to the student's:

At=log⁡pTd(yt∣ht)−log⁡πu(yt∣ht)A_t = \log p_{T_d}(y_t \mid h_t) - \log \pi_u(y_t \mid h_t)At​=logpTd​​(yt​∣ht​)−logπu​(yt​∣ht​)

where ht=(x,y<t)h_t = (x, y_{<t})ht​=(x,y<t​). A positive AtA_tAt​ encourages the sampled token, while a negative value discourages it, with the magnitude weighting the policy-gradient contribution.

However, independently trained specialists produce advantages on different scales. For instance, IF log-ratios can be significantly more dispersed than mathematics log-ratios. To address this, DN-MOPD applies domain normalization using the standard deviation of the feedback within each domain. For a response iii, the log-ratio is defined as ri,t=ℓi,tT−ℓi,trollr_{i,t} = \ell_{i,t}^T - \ell_{i,t}^{\text{roll}}ri,t​=ℓi,tT​−ℓi,troll​, where ℓT\ell^TℓT is the teacher's token log-probability and ℓroll\ell^{\text{roll}}ℓroll is the student's cached rollout log-probability. Within each batch, σd\sigma_dσd​ represents the population standard deviation of ri,tr_{i,t}ri,t​ over valid tokens for domain ddd, and σall\sigma_{\text{all}}σall​ is computed over all valid tokens pooled across domains. Each domain is assigned a multiplier:

wd=clip(σallσd,0.25,4),A~t=wdAtw_d = \text{clip}\left(\frac{\sigma_{\text{all}}}{\sigma_d}, 0.25, 4\right), \quad \widetilde{A}_t = w_d A_twd​=clip(σd​σall​​,0.25,4),At​=wd​At​

This rule amplifies feedback with a smaller spread and attenuates feedback with a larger spread, targeting dispersion rather than treating the statistic as an estimate of teacher quality. By multiplying by a positive factor without subtracting the domain mean, the method preserves the sign of each advantage. The bounds [0.25,4][0.25, 4][0.25,4] limit the adjustment when scale ratios are extreme.

The complete training iteration is illustrated below.

For the student update, the actor recomputes the pre-update student log-probabilities ℓactor\ell^{\text{actor}}ℓactor on the sampled responses to form Ai,t=ℓi,tT−ℓi,tactorA_{i,t} = \ell_{i,t}^T - \ell_{i,t}^{\text{actor}}Ai,t​=ℓi,tT​−ℓi,tactor​. The scaled advantage is detached, A~i,t=stopgrad(wdiAi,t)\widetilde{A}_{i,t} = \text{stopgrad}(w_{d_i} A_{i,t})Ai,t​=stopgrad(wdi​​Ai,t​), and utilized in the baseline clipped OPD loss. Defining the policy ratio as ρi,t(θ)=πθ(yi,t∣hi,t)/exp⁡(ℓi,tactor)\rho_{i,t}(\theta) = \pi_\theta(y_{i,t} \mid h_{i,t}) / \exp(\ell_{i,t}^{\text{actor}})ρi,t​(θ)=πθ​(yi,t​∣hi,t​)/exp(ℓi,tactor​), the token loss is computed as:

Li,t(θ)=−min⁡{ρi,t(θ)A~i,t,clip(ρi,t(θ),1−η−,1+η+)A~i,t}\mathcal{L}_{i,t}(\theta) = -\min\left\{\rho_{i,t}(\theta)\widetilde{A}_{i,t}, \text{clip}\left(\rho_{i,t}(\theta), 1 - \eta_-, 1 + \eta_+\right)\widetilde{A}_{i,t}\right\}Li,t​(θ)=−min{ρi,t​(θ)Ai,t​,clip(ρi,t​(θ),1−η−​,1+η+​)Ai,t​}

where η−\eta_-η−​ and η+\eta_+η+​ are the policy-ratio clipping bounds. The loss reduction averages valid token losses within each response and then averages across responses. Setting every wdw_dwd​ to one recovers the standard label-routed OPD baseline.

Experiment

The experiments evaluate DN-MOPD against label-routed OPD and multiple baselines using Qwen3.5 models at 9B, 4B, and 2B on mathematics, code generation, and instruction following tasks, with sensitivity checks on evaluation length and training duration. DN-MOPD consistently outperforms label-routed OPD across model sizes and evaluation budgets, with gains concentrated in mathematics and driven by recalibrating per-domain feedback weights rather than changing teacher assignment or using longer responses. Ablations indicate the improvement comes mainly from amplifying mathematics feedback and downweighting instruction-following feedback, which otherwise dominates the shared update despite contributing few tokens. The advantage also holds over the strongest single-teacher student and persists under longer training and evaluation budgets, although the eventual ordering at convergence remains unresolved.

At 9B, multi-teacher integration shows DN-MOPD improving over the label-routed baseline mainly through stronger mathematics transfer. The strongest single-teacher student remains a competitive reference, and the label-routed baseline does not surpass it, while DN-MOPD does. Gains in code and instruction following are smaller and less consistent, indicating the clearest benefit is from improving how label-routed feedback is used. DN-MOPD outperforms the strongest single-teacher student at 9B, and its advantage widens with extended training. Mathematics contributes most of the improvement, while code and instruction-following gains remain smaller and less consistent. The label-routed baseline never exceeds the strongest single-teacher student, whereas DN-MOPD does.

At smaller scales, individual specialists improve their own domains over the initial student, but multi-teacher integration gains are uneven. DN-MOPD improves over the label-routed baseline at every model size, with mathematics as the main source of gain while code and instruction following are smaller or less consistent. The benefit is attributed primarily to calibrating feedback scale, especially reducing instruction-following influence, rather than changing teacher selection or using longer generations. Individual experts improve their specialty domains at both smaller scales, while their effects on other domains vary. DN-MOPD exceeds the label-routed baseline across sizes, with mathematics driving the clearest gains and instruction-following downweighting contributing most of the improvement.

Relative to label-routed OPD, DN-MOPD improves total score across all tested model sizes, while changing teacher assignment through uniform pooling or dynamic routing produces inconsistent and sometimes negative changes. A fixed reduction of the IF weight to 0.25 recovers most of the DN-MOPD gain at 4B and 2B, whereas doubling the math weight gives only partial or little benefit. This supports attributing the gains mainly to feedback-scale calibration, particularly limiting IF influence, rather than to teacher-selection changes. DN-MOPD is the only variant that improves over Label at every tested model size, with larger gains at 4B and 2B than at 9B. Uniform pooling and dynamic routing bring mixed or negative changes across sizes, unlike DN-MOPD's consistent improvement. Downweighting IF to 0.25 alone nearly matches DN-MOPD at 4B and 2B, while doubling math weight recovers only part of the gain at 4B and little at 2B.

The data reports mean response tokens by domain and the share of outputs reaching the 16K generation cap for several 9B methods and one 4B baseline. Among 9B methods, Label-routed has the longest mathematics responses and the highest cap rate, while DN-MOPD lowers both, consistent with its mathematics gains coming from better feedback calibration rather than longer generations. Instruction-following tokens remain a small fraction of response volume across methods, even though the accompanying analysis describes their training signal as disproportionately influential. Across the listed 9B methods, Label-routed generates the longest mathematics responses and the largest share of outputs hitting the 16K cap, while DN-MOPD reduces both measures. SeqKD-SFT and ParamMerge-TA show the lowest cap rates among the 9B methods, suggesting their outputs are less likely to run into the generation limit. Instruction-following responses account for only a small fraction of domain tokens in every listed method, aligning with the observation that instruction-following supplies about 1% of response tokens but can dominate gradient variance. The 4B SeqKD-SFT baseline has longer code responses than mathematics responses, whereas the 9B SeqKD-SFT model has more similar mathematics and code lengths.

Extending training from 80 to 160 updates produces modest gains in six-task total at the 16K evaluation setting. DN-MOPD remains ahead of Label-routed and single-teacher baselines at every model size and both training endpoints, although longer training narrows the gaps at smaller sizes. Longer training improves reported methods, with single-teacher models showing smaller gains than Label-routed and DN-MOPD. DN-MOPD maintains the highest six-task total at both 80 and 160 updates and across model sizes. The advantage persists beyond the main training endpoint, even as longer runs reduce the margin at smaller sizes.

The experiments compare DN-MOPD with label-routed OPD and single-teacher baselines across 9B, 4B, and 2B models on six tasks spanning mathematics, code, and instruction following. DN-MOPD consistently improves over the label-routed baseline at every scale, outperforms the strongest single-teacher model at 9B, and shows the clearest gains in mathematics while code and instruction-following improvements are smaller or inconsistent. Ablations attribute the gain mainly to feedback-scale calibration, especially downweighting instruction-following influence, rather than to teacher selection or longer generations; response-length results corroborate this by showing reduced math response length and cap rate. Extended training keeps DN-MOPD ahead of baselines at all tested sizes, though longer runs narrow the margin at smaller scales.


KI mit KI entwickeln

Von der Idee bis zum Launch – beschleunigen Sie Ihre KI-Entwicklung mit kostenlosem KI-Co-Coding, sofort einsatzbereiter Umgebung und bestem GPU-Preis.

KI-gestütztes kollaboratives Programmieren
Sofort einsatzbereite GPUs
Die besten Preise

HyperAI Newsletters

Abonnieren Sie unsere neuesten Updates
Wir werden die neuesten Updates der Woche in Ihren Posteingang liefern um neun Uhr jeden Montagmorgen
Unterstützt von MailChimp