Command Palette
Search for a command to run...
IDバランシング:PIDベースの負荷制御による極めてスパースなMoEの安定な学習
IDバランシング:PIDベースの負荷制御による極めてスパースなMoEの安定な学習
概要
Mixture-of-Experts(MoE)による大規模言語モデル(LLM)のスケーリングは、トークンあたりの計算量をほぼ一定に保ったまま、パラメータ数を大幅に増やすことを可能にします。しかし、パラメータ数をさらに増やすには、ますますスパースなルーティングが必要となり、エキスパートの負荷不均衡が深刻化します。この不均衡は、パラメータの利用率と学習効率を低下させ、学習の安定性を損ない、信頼性の高いスケーリングのボトルネックとなる可能性があります。本研究では、補助損失を用いない代表的な2つの手法を、不完全な比例・積分・微分(PID)コントローラとして統一的に捉えます。DeepSeekの損失フリー手法は固定ステップの積分コントローラとして機能し、Kimi K3のQuantile Balancingは一般化された比例コントローラとして機能します。この制御の視点に基づき、積分・微分(ID)コントローラであるIDバランシングを提案します。IDバランシングは、積分項を負荷誤差に応じてスケーリングし、微分項は不均衡が悪化した場合にのみ活性化することで、大きな誤差や悪化する誤差に対してはより強力な補正を行い、平衡に近い状態では小さな更新を行います。768エキスパートにわたるTop-10、Top-5、Top-3ルーティングでの評価において、IDバランシングは、Top-3設定で最良のベースラインと比較して、最悪ケースのバックボーンMaxVioと学習平均のバックボーンMinVioをそれぞれ50%以上および12%削減します。総パラメータ数が18.9Bから69.9B(Top-10-of-768)に増加した場合、IDバランシングの最悪ケースのバックボーンMaxVioはほぼ変化せず、補助損失ベースラインよりも約89.6%低くなります。また、IDバランシングは、競争力のある言語モデリングおよび下流タスクのパフォーマンスも維持します。IDバランシングの利点はスパース性が高まるにつれて拡大し、より大規模でスパースなMoEモデルをスケーリングするための有望なソリューションとなります。
One-sentence Summary
The Qwen Team and Alibaba Group propose ID Balancing, a PID-inspired Integral–Derivative load controller for extremely sparse MoE training that unifies auxiliary-loss-free methods as incomplete PID controllers, scales its integral term with load error, and activates its derivative term only when imbalance worsens, cutting worst-case backbone MaxVio by over 50% and 12% relative to the best Top-3 baselines while remaining stable as parameters scale from 18.9B to 69.9B.
Key Contributions
- Introduces a unified PID control perspective for auxiliary-loss-free MoE balancing, framing DeepSeek’s loss-free method as fixed-step integral control and Kimi K3’s Quantile Balancing as generalized proportional control.
- Presents ID Balancing, an Integral–Derivative controller that scales its integral term with load error and activates a derivative term only when imbalance worsens, using O(E) token-count feedback per layer without adding a balancing gradient to the language-modeling objective.
- Across Top-10, Top-5, and Top-3 routing over 768 experts, ID Balancing reduces worst-case backbone MaxVio by over 50% and training-average backbone MinVio by over 12% relative to best baselines in the Top-3 setting, keeps worst-case backbone MaxVio nearly unchanged when scaling from 18.9B to 69.9B parameters, and maintains competitive language-modeling and downstream performance.
Introduction
Mixture-of-Experts (MoE) models scale large language models efficiently through sparse expert activation, where a router selects only K experts per token from a larger pool, allowing capacity growth without proportional increases in computation. As the expert pool expands, maintaining balanced expert loads becomes increasingly difficult; small shifts near the Top-K selection boundary can cause overloaded experts to slow computation and underloaded experts to remain undertrained, limiting both training efficiency and model capacity benefits.
Existing auxiliary-loss-free balancing methods adjust expert biases based on routing feedback, but they operate with fixed-step corrections or batch-local targets that do not adapt to error magnitude or worsening imbalance. The authors interpret these methods through a unified Proportional-Integral-Derivative (PID) control framework, identifying DeepSeek's loss-free method as fixed-step integral control and Kimi K3's Quantile Balancing as generalized proportional control. This perspective reveals limitations in handling highly sparse routing and motivates the authors' main contribution: ID Balancing, an Integral-Derivative controller that applies magnitude-aware integral corrections and a worsening-gated derivative term, requiring only O(E) token-count feedback per layer without adding a balancing gradient to the training objective.
Method
The authors formulate Mixture-of-Experts (MoE) load balancing as a control problem to manage expert utilization during training. In this framework, an MoE layer with E experts uses a Top-K router that computes logits z=Wrx for a token representation x. A non-trainable bias vector b is added to the expert scores before selection, such that the selected experts are T(x)=TopKi∈[E](si+bi). The control target is a uniform load across experts. At each training step t, the normalized load error for expert i is defined as ei(t)=(nˉ(t)−ni(t))/nˉ(t), where nˉ(t) is the average token count and ni(t) is the count for expert i. The bias vector serves as both the persistent controller state and the routing input.
To interpret existing loss-free balancing methods, the authors adopt a generalized Proportional-Integral-Derivative (PID) control view, where the bias update is expressed as b(t+1)=P(t)+I(t)+D(t). Under this lens, DeepSeek's loss-free method operates as a fixed-step integral controller that updates the bias based solely on the sign of the load error, while Quantile Balancing functions as a generalized proportional controller that sets the bias directly from a target computed on the current batch score distribution.
Building on this control perspective, the authors propose ID Balancing, which combines a magnitude-aware integral term with a worsening-gated derivative term, while omitting the proportional term to maintain stable routing boundaries.
The magnitude-aware integral term addresses the limitation of fixed-step integral control, which applies the same correction regardless of error severity. Instead, ID Balancing scales the update by the normalized error itself:
bi(t+1)=bi(t)+Kiei(t)where Ki≥0 is the integral gain. This ensures that the correction grows with the magnitude of the imbalance and shrinks as the expert approaches its target load. Because the sum of normalized errors is zero, this integral update inherently preserves a zero-mean bias when initialized at zero.
As shown in the figure above, scaling the integral update by the error magnitude significantly reduces the worst-layer maximum violation (MaxVio) peak and lowers sustained underload compared to the fixed-sign update used in DeepSeek's method. This demonstrates stronger early correction and better expert utilization.
To further capture the evolution of load errors, ID Balancing introduces a worsening-gated derivative term. This term computes the change in error Δei(t)=ei(t)−ei(t−1) and applies a gate gi(t) that activates only when the imbalance worsens. The gate opens when the previous error and its change share the same sign, indicating that the deviation is growing without crossing zero. The derivative correction Kdgi(t)Δei(t) is accumulated in the bias alongside the integral update.
Refer to the framework diagram above, which illustrates the behavior of the derivative term. The active gate fraction decreases as training progresses, and adding this gated correction lowers early load violations and reduces expert concentration, while the trajectories eventually approach those of the integral-only update.
The authors deliberately omit the proportional term in ID Balancing. Proportional controllers like Quantile Balancing compute a target bias from the current batch, which changes with the score distribution and leads to larger late-stage bias drift. Smaller changes in relative expert biases limit perturbations to the Top-K selection boundary, stabilizing routing and facilitating weight merging.
As demonstrated in the figure above, while exponential moving average (EMA) smoothing can reduce bias drift in Quantile Balancing, it weakens early load control. ID Balancing achieves low bias drift and effective load balance through its integral and gated derivative updates without requiring additional smoothing.
Finally, because the gating mechanism in the derivative term does not preserve the zero-mean property, ID Balancing includes a zero-mean centering step. After computing the intermediate bias b~i(t+1), the common component is removed by subtracting the mean across all experts:
bi(t+1)=b~i(t+1)−E1j=1∑Eb~j(t+1)This ensures that the sum of the biases remains zero, preventing common bias drift without altering the relative score ranking or the selected expert set.
The complete ID Balancing update effectively limits transient overload and sustained underload in highly sparse MoE training. As shown in the figure above, ID Balancing maintains lower mean MaxVio and Mean MinVio throughout training compared to auxiliary loss, DeepSeek's loss-free method, and Quantile Balancing, ensuring stable and efficient expert utilization.
Experiment
The experiments validate ID Balancing across standard and continued pretraining, higher learning rates, and inference-time utilization. ID Balancing consistently improves backbone load balance over Auxiliary loss and DeepSeek's loss-free method while maintaining competitive LM loss and downstream quality, with reduced gains preserving performance during continued pretraining. At higher learning rates, it sustains effective load control with smoother gradient norms, though Quantile Balancing achieves tighter backbone balance in this stress test. Ablations confirm the default settings of K_i = K_d = 6e-3, balancing early correction, worst-case violations, and MTP trade-offs.
ID Balancing achieves the lowest backbone load imbalance metrics among the compared methods while keeping language modeling loss competitive, and it also improves expert utilization at inference. The comparisons are made under Top-3, Top-5, and Top-10 routing over 768 experts in 18.9B parameter models trained on 120B tokens. ID Balancing consistently outperforms Auxiliary loss, DeepSeek's loss-free method, and Quantile Balancing in reducing backbone MaxVio and MinVio, with particularly notable gains in worst-case and final-step load control. ID Balancing yields the lowest backbone MaxVio across last-1k, average, and worst-case metrics compared to all other methods, with the worst-case being several times smaller than Auxiliary loss. ID Balancing reduces the average inactive-expert ratio under Top-10-of-768 routing from 8.4% (Auxiliary loss) to 6.4%, and the first-layer ratio from 9.3% to 2.6%. ID Balancing maintains competitive language modeling loss, with values within a narrow range of the best-performing method, while providing substantially better load balance. Quantile Balancing and ID Balancing both show more consistent load control across layers than Auxiliary loss or DeepSeek's loss-free method, with smaller differences between training-average and final-stage profiles. MinVio for ID Balancing is lower than or comparable to other methods, indicating better expert utilization with fewer underutilized experts.
ID Balancing maintains competitive downstream performance compared to other routing methods across nine benchmarks, with strong results in knowledge-heavy tasks and code generation. Its average score is among the best in both the Top-8-of-256 and Top-10-of-768 configurations, often matching or exceeding baselines like Auxiliary loss and DeepSeek loss-free. ID Balancing achieves the highest average score in both routing configurations, edging out Auxiliary loss and DeepSeek loss-free. ID Balancing leads on MMLU-Pro and SuperGPQA in the Top-8 configuration, indicating robust knowledge and reasoning capabilities. Under Top-10-of-768 routing, ID Balancing posts the best or near-best scores on most individual benchmarks, including strong code generation results.
At a constant learning rate 2.3 times the standard peak, both ID Balancing and Quantile Balancing maintain lower backbone load violations than Auxiliary loss and DeepSeek's loss-free method. Quantile Balancing achieves the tightest backbone balance, while ID Balancing yields the lowest LM loss and the lowest last-1k MTP load violations, indicating different trade-offs between backbone and MTP balance. ID Balancing and Quantile Balancing produce backbone MaxVio values around 0.64 and 0.50, compared to 3.14 and 1.38 for the two baselines. Quantile Balancing achieves the lowest backbone last-1k MaxVio among all methods. ID Balancing achieves the lowest LM loss (2.0007) and the lowest last-1k MTP MaxVio. Auxiliary loss shows a very high worst-case backbone MaxVio of 20.89, while Quantile Balancing's is only 2.96.
The integral gain sweep shows that increasing the gain from 3 to 6 reduces backbone violations, but a further increase to 9 worsens worst-case MTP balance and backbone worst-case values. The intermediate gain of 6 is chosen as it balances early correction with lower worst-case violations. The derivative gain is then evaluated separately, showing early backbone load improvements come at the cost of MTP balance. Raising the integral gain from 3 to 6 improves both average and worst-case backbone violations, while a further increase to 9 only slightly improves the average but significantly worsens worst-case backbone and MTP violations. The smallest integral gain yields the lowest LM loss and better final-stage and MTP metrics, but corrects early imbalance more slowly. Adding a derivative term improves early backbone load control, reducing backbone violations over early steps, but increases MTP violations, with larger derivative gains amplifying this trade-off.
Increasing the derivative gain from zero to 12e-3 progressively improves early backbone load balancing, reducing backbone overload metrics over the first 5k steps. However, this comes at the cost of higher MTP module overload, particularly in later and average metrics. The default gain of 6e-3 offers a balance, yielding the lowest LM loss and training-average backbone MinVio among tested values. Adding a derivative term to the integral-only baseline reduces backbone MaxVio over both 0–1k and 1–5k steps, with larger gains giving stronger early corrections. The default derivative gain of 6e-3 lowers training-average backbone MaxVio from 0.8003 to 0.7931 relative to the integral-only baseline. MTP overload increases with larger derivative gains; training-average MTP MaxVio rises from 0.8862 (baseline) to 1.0922 at the largest gain. The default gain achieves the lowest LM loss (1.7426) and the lowest training-average backbone MinVio among the compared settings.
ID Balancing consistently improves expert utilization and backbone load balance across Top-3, Top-5, and Top-10 routing over 768 experts in 18.9B models, reducing inactive-expert ratios and worst-case load violations while keeping language modeling loss competitive. It also matches or exceeds baselines on nine downstream benchmarks, especially in knowledge-heavy tasks and code generation. A constant high learning rate shows ID Balancing and Quantile Balancing maintain lower backbone violations than the baselines, but with different trade-offs: Quantile Balancing achieves tighter backbone balance, while ID Balancing yields lower LM loss and MTP violations. Gain sweeps reveal that moderate integral and derivative gains improve early load correction without sacrificing worst-case balance, though larger derivative gains worsen MTP overload.