Command Palette
Search for a command to run...
教師割り当てを超えて:ドメイン正規化マルチ教師オンポリシー蒸留
教師割り当てを超えて:ドメイン正規化マルチ教師オンポリシー蒸留
Xin Li Hao Jiang Xin Gao Annan Wang Yuchen Xie Jinghao Guo Xingwei Qu Yichi Zhang Chau Yuen
概要
強化学習は、一つの言語モデルを数学・コーディング・指示追従といった単一スキルに優れた複数の専門家モデルへと変えることができるが、ユーザーが必要とするのはこれらのスキルをすべて備えた一つのモデルである。マルチ教師オンポリシー蒸留(MOPD)は、専門家たちに一人の生徒モデルを教えさせることでこれらを統合する。すなわち、生徒モデルが各プロンプトに回答し、そのプロンプトのドメインを担当する専門家がすべてのトークンに対してフィードバックを与える。このルーティングはどの専門家が教えるかを決定するが、そのフィードバックが共有の生徒モデルをどの程度強く更新するかは決定しない。3つの規模のQwen3.5モデルにおいて、MOPDの生徒モデルは最良の単一専門家に教えられたモデルを上回らず、数学専門家の利点をほとんど獲得しないことが分かった。フィードバックには不均衡があり、指示追従のフィードバックは数学のフィードバックよりも数倍ばらつきが大きく、生徒モデルの更新を支配する。我々はドメイン正規化MOPD(DN-MOPD)を提案する。これはルーティングを維持しつつ、各ドメインのフィードバックをその測定されたばらつきによって再スケーリングする。6つの公開ベンチマークにおいて、DN-MOPDは、すべてのモデル規模、3つのランダムシード、2つの回答長制限の下で、MOPDを上回る平均スコアを達成し、失われた数学の利得の大部分を回復する。固定ドメイン重みを用いた対照実験により、この利得は主に数学フィードバックだけを強めることではなく、指示追従フィードバックを弱めることに由来すること、またDN-MOPDが測定する重みに近い固定重みでも同程度の性能が得られることが示される。したがって、専門家の統合には、どの専門家が教えるかだけでなく、そのフィードバックをどの程度強く反映するかも決定することが必要である。
One-sentence Summary
Researchers from Nanyang Technological University, Yale University, and University of Manchester propose Domain-Normalized Multi-Teacher On-Policy Distillation (DN-MOPD), which keeps specialist routing but rescales each domain's feedback by its measured spread to correct unbalanced feedback, and on six public benchmarks DN-MOPD outperforms MOPD across model sizes, seeds, and answer-length limits while recovering most of the lost mathematics gains.
Key Contributions
- In Qwen3.5 models at three sizes, multi-teacher on-policy distillation (MOPD) exhibits unbalanced feedback: instruction-following feedback has several times the spread of mathematics feedback, so the student gains little from the mathematics specialist and does not beat the strongest single-teacher baseline.
- Domain-Normalized MOPD (DN-MOPD) keeps MOPD's teacher routing and rescales each domain's feedback by the measured spread of its teacher-student log-ratios.
- On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and two answer-length limits, and recovers most of the lost mathematics gain; fixed-weight controls show the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone.
Introduction
Post-training for large language models increasingly uses domain-specific reinforcement learning, producing separate experts for mathematics, code, and instruction following that must be combined into one deployable model. Multi-teacher on-policy distillation (MOPD) integrates these experts by routing each prompt to its domain expert and using that expert’s token-level log-probabilities on the student’s own response, but routing alone does not control how much each expert’s feedback counts. Fixed or gap-based domain weights ignore differences in feedback spread, so dispersed instruction-following feedback can dominate the shared gradient while mathematics transfer remains weak. The authors introduce Domain-Normalized Multi-Teacher On-Policy Distillation (DN-MOPD), which estimates the spread of teacher–student log-ratios within each domain on every batch and applies a bounded multiplier to bring each domain’s distillation advantages toward the pooled scale. Across Qwen3.5 expert pools at 9B, 4B, and 2B, DN-MOPD improves the six-task average over MOPD and recovers much of the mathematics gain that MOPD loses, without adding extra teacher models, teacher calls, or learned routers.
Method
The authors propose DN-MOPD, a method that retains the domain-label assignment of multi-teacher Online Policy Distillation (OPD) while introducing a domain normalization step. Before the shared student model is updated, the distillation advantages for each domain are rescaled based on their feedback spread relative to the current batch.
The training setup begins with one initial model. Three copies are trained independently via reinforcement learning to become specialists for mathematics, code, and instruction following (IF). A fourth copy serves as the student πu, which is trained to acquire all three skills. Each training prompt carries a domain label d∈{math, code, IF}, and Td denotes the corresponding domain specialist.
To generate feedback, the student first produces a response y=(y1,…,yn) to a prompt. The matching domain teacher then evaluates this response, calculating the token-level distillation advantage by comparing the teacher's token probability to the student's:
At=logpTd(yt∣ht)−logπu(yt∣ht)where ht=(x,y<t). A positive At encourages the sampled token, while a negative value discourages it, with the magnitude weighting the policy-gradient contribution.
However, independently trained specialists produce advantages on different scales. For instance, IF log-ratios can be significantly more dispersed than mathematics log-ratios. To address this, DN-MOPD applies domain normalization using the standard deviation of the feedback within each domain. For a response i, the log-ratio is defined as ri,t=ℓi,tT−ℓi,troll, where ℓT is the teacher's token log-probability and ℓroll is the student's cached rollout log-probability. Within each batch, σd represents the population standard deviation of ri,t over valid tokens for domain d, and σall is computed over all valid tokens pooled across domains. Each domain is assigned a multiplier:
wd=clip(σdσall,0.25,4),At=wdAtThis rule amplifies feedback with a smaller spread and attenuates feedback with a larger spread, targeting dispersion rather than treating the statistic as an estimate of teacher quality. By multiplying by a positive factor without subtracting the domain mean, the method preserves the sign of each advantage. The bounds [0.25,4] limit the adjustment when scale ratios are extreme.
The complete training iteration is illustrated below.
For the student update, the actor recomputes the pre-update student log-probabilities ℓactor on the sampled responses to form Ai,t=ℓi,tT−ℓi,tactor. The scaled advantage is detached, Ai,t=stopgrad(wdiAi,t), and utilized in the baseline clipped OPD loss. Defining the policy ratio as ρi,t(θ)=πθ(yi,t∣hi,t)/exp(ℓi,tactor), the token loss is computed as:
Li,t(θ)=−min{ρi,t(θ)Ai,t,clip(ρi,t(θ),1−η−,1+η+)Ai,t}where η− and η+ are the policy-ratio clipping bounds. The loss reduction averages valid token losses within each response and then averages across responses. Setting every wd to one recovers the standard label-routed OPD baseline.
Experiment
The experiments evaluate DN-MOPD against label-routed OPD and multiple baselines using Qwen3.5 models at 9B, 4B, and 2B on mathematics, code generation, and instruction following tasks, with sensitivity checks on evaluation length and training duration. DN-MOPD consistently outperforms label-routed OPD across model sizes and evaluation budgets, with gains concentrated in mathematics and driven by recalibrating per-domain feedback weights rather than changing teacher assignment or using longer responses. Ablations indicate the improvement comes mainly from amplifying mathematics feedback and downweighting instruction-following feedback, which otherwise dominates the shared update despite contributing few tokens. The advantage also holds over the strongest single-teacher student and persists under longer training and evaluation budgets, although the eventual ordering at convergence remains unresolved.
At 9B, multi-teacher integration shows DN-MOPD improving over the label-routed baseline mainly through stronger mathematics transfer. The strongest single-teacher student remains a competitive reference, and the label-routed baseline does not surpass it, while DN-MOPD does. Gains in code and instruction following are smaller and less consistent, indicating the clearest benefit is from improving how label-routed feedback is used. DN-MOPD outperforms the strongest single-teacher student at 9B, and its advantage widens with extended training. Mathematics contributes most of the improvement, while code and instruction-following gains remain smaller and less consistent. The label-routed baseline never exceeds the strongest single-teacher student, whereas DN-MOPD does.
At smaller scales, individual specialists improve their own domains over the initial student, but multi-teacher integration gains are uneven. DN-MOPD improves over the label-routed baseline at every model size, with mathematics as the main source of gain while code and instruction following are smaller or less consistent. The benefit is attributed primarily to calibrating feedback scale, especially reducing instruction-following influence, rather than changing teacher selection or using longer generations. Individual experts improve their specialty domains at both smaller scales, while their effects on other domains vary. DN-MOPD exceeds the label-routed baseline across sizes, with mathematics driving the clearest gains and instruction-following downweighting contributing most of the improvement.
Relative to label-routed OPD, DN-MOPD improves total score across all tested model sizes, while changing teacher assignment through uniform pooling or dynamic routing produces inconsistent and sometimes negative changes. A fixed reduction of the IF weight to 0.25 recovers most of the DN-MOPD gain at 4B and 2B, whereas doubling the math weight gives only partial or little benefit. This supports attributing the gains mainly to feedback-scale calibration, particularly limiting IF influence, rather than to teacher-selection changes. DN-MOPD is the only variant that improves over Label at every tested model size, with larger gains at 4B and 2B than at 9B. Uniform pooling and dynamic routing bring mixed or negative changes across sizes, unlike DN-MOPD's consistent improvement. Downweighting IF to 0.25 alone nearly matches DN-MOPD at 4B and 2B, while doubling math weight recovers only part of the gain at 4B and little at 2B.
The data reports mean response tokens by domain and the share of outputs reaching the 16K generation cap for several 9B methods and one 4B baseline. Among 9B methods, Label-routed has the longest mathematics responses and the highest cap rate, while DN-MOPD lowers both, consistent with its mathematics gains coming from better feedback calibration rather than longer generations. Instruction-following tokens remain a small fraction of response volume across methods, even though the accompanying analysis describes their training signal as disproportionately influential. Across the listed 9B methods, Label-routed generates the longest mathematics responses and the largest share of outputs hitting the 16K cap, while DN-MOPD reduces both measures. SeqKD-SFT and ParamMerge-TA show the lowest cap rates among the 9B methods, suggesting their outputs are less likely to run into the generation limit. Instruction-following responses account for only a small fraction of domain tokens in every listed method, aligning with the observation that instruction-following supplies about 1% of response tokens but can dominate gradient variance. The 4B SeqKD-SFT baseline has longer code responses than mathematics responses, whereas the 9B SeqKD-SFT model has more similar mathematics and code lengths.
Extending training from 80 to 160 updates produces modest gains in six-task total at the 16K evaluation setting. DN-MOPD remains ahead of Label-routed and single-teacher baselines at every model size and both training endpoints, although longer training narrows the gaps at smaller sizes. Longer training improves reported methods, with single-teacher models showing smaller gains than Label-routed and DN-MOPD. DN-MOPD maintains the highest six-task total at both 80 and 160 updates and across model sizes. The advantage persists beyond the main training endpoint, even as longer runs reduce the margin at smaller sizes.
The experiments compare DN-MOPD with label-routed OPD and single-teacher baselines across 9B, 4B, and 2B models on six tasks spanning mathematics, code, and instruction following. DN-MOPD consistently improves over the label-routed baseline at every scale, outperforms the strongest single-teacher model at 9B, and shows the clearest gains in mathematics while code and instruction-following improvements are smaller or inconsistent. Ablations attribute the gain mainly to feedback-scale calibration, especially downweighting instruction-following influence, rather than to teacher selection or longer generations; response-length results corroborate this by showing reduced math response length and cap rate. Extended training keeps DN-MOPD ahead of baselines at all tested sizes, though longer runs narrow the margin at smaller scales.