Command Palette
Search for a command to run...
Repenser la distillation on-policy entre tokeniseurs : de la couverture d’alignement à la fiabilité de la supervision
Repenser la distillation on-policy entre tokeniseurs : de la couverture d’alignement à la fiabilité de la supervision
Bingxi Hou Guochao Jiang Guofeng Quan Weiqing Li Wenfeng Feng Guohua Liu Yuewei Zhang
Résumé
La distillation on-policy (OPD) entraîne un modèle étudiant sur ses propres générations en utilisant le retour de l’enseignant. Lorsque les tokeniseurs diffèrent, comparer les prédictions de l’enseignant et de l’étudiant nécessite un alignement à la fois au niveau des séquences et du vocabulaire. Dans cet article, nous examinons si l’élargissement de cette couverture d’alignement améliore l’apprentissage. Sur trois paires enseignant–étudiant hétérogènes portant sur le raisonnement mathématique et la génération de code, les groupes stricts 1:1 couvrent déjà la plupart des jetons générés par l’étudiant malgré une divergence substantielle de vocabulaire. Sur des réponses échantillonnées auprès des étudiants avant distillation, le vocabulaire partagé conserve en moyenne presque toute la masse de probabilité de l’enseignant et de l’étudiant aux positions strictement alignées. Restreindre la KL inverse à un sous-ensemble des 16 premiers jetons du vocabulaire partagé sélectionnés par l’étudiant à chaque position stricte permet d’obtenir une précision comparable à celle d’une OPD sur tout le vocabulaire partagé, en surpassant les références inter-tokeniseurs évaluées. L’ajout d’une supervision par erreur quadratique moyenne sur les log-probabilités des segments dans les groupes de non-correspondance offre une couverture de supervision complète, mais réduit la précision. Aux points de contrôle issus d’un entraînement avec uniquement la perte stricte, les gradients des segments présentent un accord directionnel faible ou négatif avec les gradients stricts et leur amplitude augmente par rapport à ces derniers. Ces diagnostics peuvent aider à expliquer la baisse de précision due à l’ajout de la supervision sur les segments. Nos résultats incitent à passer d’une maximisation de la couverture d’alignement à une priorisation de la fiabilité de la supervision : une supervision compacte aux positions strictes peut être plus efficace qu’une couverture plus large qui introduit des signaux d’entraînement faiblement alignés ou conflictuels.
One-sentence Summary
Researchers at Alibaba Cloud Computing show that in cross-tokenizer on-policy distillation, strict 1:1 aligned groups already cover most student-generated tokens, and restricting reverse KL to a student-selected top-16 shared-vocabulary subset at strict positions matches full shared-vocabulary OPD while outperforming baselines, whereas adding span MSE supervision degrades accuracy and motivates a shift from maximizing alignment coverage to prioritizing supervision reliability.
Key Contributions
- The paper analyzes alignment coverage in cross-tokenizer on-policy distillation across mathematical reasoning and code generation tasks, showing that strict 1:1 aligned groups cover most student-generated tokens and retain nearly all teacher and student probability mass at strict positions despite vocabulary mismatch.
- It demonstrates that restricting reverse KL distillation to a student-selected top-16 subset of the shared vocabulary at strict positions achieves accuracy comparable to full shared-vocabulary distillation while outperforming the evaluated cross-tokenizer baselines.
- The work shows that adding span log-probability MSE supervision on mismatch groups reduces downstream accuracy, and its gradient diagnostics reveal weak or negative agreement with strict gradients along with growing relative magnitude, motivating a shift from maximizing alignment coverage to prioritizing supervision reliability.
Introduction
On-policy distillation trains a language model on its own generated responses using teacher feedback, which matters because it aligns the student with teacher preferences at the states the student actually visits. When the student and teacher use different tokenizers, the same response can have different token boundaries and next-token distributions over different vocabularies, so cross-tokenizer distillation must handle both sequence-level and vocabulary-level alignment. Prior methods such as rank matching, learned mappings, likelihood matching, byte-level outputs, and multi-token grouping have aimed to recover more supervision, but they generally emphasize alignment coverage without fully testing whether the added supervision is useful for learning. The authors examine strict cross-tokenizer distillation and find that strict alignment already covers most student-generated tokens despite large static vocabulary gaps, and that a compact student-selected top-k subset of the shared vocabulary retains nearly all distillation gains. By contrast, adding span log-probability MSE for unmatched spans achieves complete coverage but reduces accuracy, motivating a shift from alignment coverage to supervision reliability.
Method
On-Policy Distillation (OPD) trains a student language model by sampling trajectories from the student itself and aligning its distribution to that of a teacher model on the prefixes the student actually visits. This process can be viewed as a dense KL-constrained reinforcement learning setup where the teacher distribution induces a token-level reward. With a shared tokenizer, OPD minimizes the following objective:
LOPD(θ)=Ex∼D,y∼πθ(⋅∣x)[i=1∑LKL(πθ(⋅∣x,y<i)∥πT(⋅∣x,y<i))].When extending this approach across model families with different tokenizers, the same text can be assigned different token boundaries and vocabulary entries. Cross-Tokenizer OPD must therefore account for alignment at both the sequence and vocabulary levels. The authors leverage token-group alignment to score and align the decoded student response using both models' tokenizers. The resulting student and teacher response-token sequences are denoted by y=(y1,…,yL) and v=(v1,…,vn). By retaining the token offsets shared by both sequences, the response is partitioned into pairs of aligned token groups:
S(y)={(Srθ,SrT)}r=1R.These groups are divided into strict 1:1 groups, where a single student token and a single teacher token span the same interval, and mismatch groups, which require multiple tokens on at least one side to construct the same span.
To align the predictions, the authors identify a shared vocabulary V∩=Vθ∩VT by matching underlying tokens. For strict 1:1 groups, the student and teacher distributions are restricted and renormalized over this shared vocabulary. The strict Cross-Tokenizer objective then sums the reverse KL divergence over strictly aligned positions:
L1:1(θ)=Ex∼D,y∼πθ(⋅∣x)r∈A1:1(y)∑KL(πˉθ(⋅∣x,y<ir)∥πˉT(⋅∣x,v<jr)).As shown in the figure below, diagnostic evaluations of this cross-tokenizer approach reveal that strict alignment remains high during training, shared vocabulary captures nearly all probability mass in both models, and excluding mismatch groups yields the best performance.
To evaluate the learning value of the remaining mismatch groups, the authors introduce span supervision. For each mismatch group, the probabilities of the observed token paths are computed for both the student and the teacher. These probabilities are matched using a mean squared error loss over the mismatch groups in a log-probability formulation:
Lspan(θ)=Ex∼D,y∼πθ(⋅∣x)r∈Amis(y)∑(logqθ(r)−logqT(r))2.The total training loss combines the strict objective and the span supervision with a weighting factor λ:
Lλ(θ)=L1:1(θ)+λLspan(θ).As shown in the figure below, the impact of mismatch weights on full-average accuracy across different teacher-student pairs demonstrates that strict supervision alone performs best.
By varying λ, the authors control the influence of the span loss while maintaining complete supervision coverage. However, empirical results indicate that adding span supervision for mismatch groups tends to reduce downstream accuracy across tested weights, confirming that strict matching retains the most useful supervision.
Experiment
The experiments study cross-tokenizer knowledge distillation on three teacher-student pairs using a shared pool of mathematics and code prompts, with downstream evaluation on math and code benchmarks. They validate that strict token alignment covers most positions on student-generated trajectories despite large static vocabulary mismatch, while adding span supervision for remaining mismatch groups consistently hurts accuracy. The shared vocabulary retains nearly all predictive probability mass, and distillation on a small student-selected top-k subset keeps most of the strict-supervision gains while outperforming four baselines. Gradient diagnostics show that span-loss gradients align poorly with strict-loss gradients and grow in relative magnitude during training, helping explain the negative effect of complete coverage.
Strict token coverage on student trajectories stayed high across the evaluated model pairs even when static vocabulary overlap differed substantially. The pair with the lowest vocabulary Jaccard overlap also had the highest strict coverage under both student and teacher tokenizations. Coverage remained stable across training windows, indicating that large static vocabulary mismatch can coexist with strict alignment at most positions. Strict student and teacher token coverage stayed high across all model pairs despite static vocabulary overlap ranging from about 39% to about 65%. Granite to Phi showed the lowest static vocabulary Jaccard overlap but the highest strict coverage under both tokenizations. Within each pair, strict coverage varied little across training windows, changing by at most 1.43 percentage points for student tokens and 3.75 for teacher tokens.
Strict full-vocabulary distillation consistently improves math and full-average accuracy over baseline alternatives and the base model across all three cross-tokenizer pairs, with the largest full-average gain for Granite-to-Qwen. Code accuracy is more mixed, but strict full remains at or near the best. A compact student-selected subset preserves nearly all of the strict full learning benefit. Strict full supervision leads over all baseline alternatives in math and full-average accuracy across every teacher-student pair. Student-selected top-k subsets retain almost all of the strict full distillation gain, with no consistent further benefit from a larger subset.
The evaluation examines cross-tokenizer distillation across several teacher-student pairs, measuring token coverage and downstream accuracy. Strict token coverage remains high across pairs even when static vocabulary overlap differs substantially, indicating that vocabulary mismatch does not prevent close sequence alignment. Strict full-vocabulary distillation consistently improves math and full-average accuracy over baselines, with the largest full-average gain for Granite-to-Qwen, while code accuracy is more mixed. A compact student-selected subset preserves nearly all of this strict distillation benefit, with no consistent advantage from using a larger subset.