HyperAIHyperAI

Command Palette

Search for a command to run...

إعادة النظر في التقطير وفق السياسة عبر المرمِّزات: من تغطية المحاذاة إلى موثوقية الإشراف

Bingxi Hou Guochao Jiang Guofeng Quan Weiqing Li Wenfeng Feng Guohua Liu Yuewei Zhang

الملخص

يدرِّب التقطير وفق السياسة (OPD) نموذجَ الطالب على توليداته الخاصة باستخدام تغذية راجعة من المعلم. عند اختلاف المرمِّزات، تتطلب مقارنة تنبؤات المعلم والطالب محاذاةً على مستوى التسلسل ومستوى المفردات معًا. ندرس في هذه الورقة ما إذا كان توسيع تغطية هذه المحاذاة يحسِّن التعلم. عبر ثلاثة أزواج غير متجانسة من المعلم والطالب في الاستدلال الرياضي وتوليد الكود، تغطي المجموعات الصارمة 1:1 معظم الرموز التي يولّدها الطالب بالفعل رغم عدم التطابق الكبير بين المفردات. على الاستجابات المسحوبة من نماذج الطلاب قبل التقطير، تحتفظ المفردات المشتركة في المتوسط بكل الكتلة الاحتمالية للمعلم والطالب تقريبًا عند المواضع المُحاذاة بدقة. يؤدي تقييد تباعد KL العكسي على مجموعة فرعية من أفضل 16 رمزًا في المفردات المشتركة يختارها الطالب عند كل موضع صارم إلى دقة مماثلة لتلك التي يحققها التقطير وفق السياسة بكامل المفردات المشتركة، متفوقًا على خطوط الأساس المُقيَّمة عبر المرمِّزات. تؤدي إضافة إشراف بمتوسط الخطأ التربيعي على لوغاريتمات احتمالات الامتدادات في مجموعات عدم التطابق إلى تغطية إشرافية كاملة، لكنها تقلل الدقة. عند نقاط التحقق من التدريب باستخدام الخسارة الصارمة فقط، تُظهر تدرجات الامتدادات اتفاقًا اتجاهيًا ضعيفًا أو سالبًا مع التدرجات الصارمة، وتتزايد مقاديرها نسبةً إليها. قد تساعد هذه المؤشرات التشخيصية في تفسير انخفاض الدقة الناتج عن إضافة إشراف الامتدادات. تدفع نتائجنا نحو التحول من تعظيم تغطية المحاذاة إلى إعطاء الأولوية لموثوقية الإشراف: فالإشراف المدمج عند المواضع الصارمة يمكن أن يكون أكثر فاعلية من التغطية الأوسع التي تُدخل إشارات تدريب ضعيفة المحاذاة أو متعارضة.

One-sentence Summary

Researchers at Alibaba Cloud Computing show that in cross-tokenizer on-policy distillation, strict 1:1 aligned groups already cover most student-generated tokens, and restricting reverse KL to a student-selected top-16 shared-vocabulary subset at strict positions matches full shared-vocabulary OPD while outperforming baselines, whereas adding span MSE supervision degrades accuracy and motivates a shift from maximizing alignment coverage to prioritizing supervision reliability.

Key Contributions

  • The paper analyzes alignment coverage in cross-tokenizer on-policy distillation across mathematical reasoning and code generation tasks, showing that strict 1:1 aligned groups cover most student-generated tokens and retain nearly all teacher and student probability mass at strict positions despite vocabulary mismatch.
  • It demonstrates that restricting reverse KL distillation to a student-selected top-16 subset of the shared vocabulary at strict positions achieves accuracy comparable to full shared-vocabulary distillation while outperforming the evaluated cross-tokenizer baselines.
  • The work shows that adding span log-probability MSE supervision on mismatch groups reduces downstream accuracy, and its gradient diagnostics reveal weak or negative agreement with strict gradients along with growing relative magnitude, motivating a shift from maximizing alignment coverage to prioritizing supervision reliability.

Introduction

On-policy distillation trains a language model on its own generated responses using teacher feedback, which matters because it aligns the student with teacher preferences at the states the student actually visits. When the student and teacher use different tokenizers, the same response can have different token boundaries and next-token distributions over different vocabularies, so cross-tokenizer distillation must handle both sequence-level and vocabulary-level alignment. Prior methods such as rank matching, learned mappings, likelihood matching, byte-level outputs, and multi-token grouping have aimed to recover more supervision, but they generally emphasize alignment coverage without fully testing whether the added supervision is useful for learning. The authors examine strict cross-tokenizer distillation and find that strict alignment already covers most student-generated tokens despite large static vocabulary gaps, and that a compact student-selected top-k subset of the shared vocabulary retains nearly all distillation gains. By contrast, adding span log-probability MSE for unmatched spans achieves complete coverage but reduces accuracy, motivating a shift from alignment coverage to supervision reliability.

Method

On-Policy Distillation (OPD) trains a student language model by sampling trajectories from the student itself and aligning its distribution to that of a teacher model on the prefixes the student actually visits. This process can be viewed as a dense KL-constrained reinforcement learning setup where the teacher distribution induces a token-level reward. With a shared tokenizer, OPD minimizes the following objective:

LOPD(θ)=Ex∼D,y∼πθ(⋅∣x)[∑i=1LKL(πθ(⋅∣x,y<i)∥πT(⋅∣x,y<i))].\mathcal{L}_{\mathrm{OPD}}(\theta) = \mathbb{E}_{x \sim \mathcal{D}, y \sim \pi_{\theta}(\cdot | x)} \left[ \sum_{i=1}^{L} \mathrm{KL}(\pi_{\theta}(\cdot | x, y_{<i}) \| \pi_{\mathrm{T}}(\cdot | x, y_{<i})) \right].LOPD​(θ)=Ex∼D,y∼πθ​(⋅∣x)​[i=1∑L​KL(πθ​(⋅∣x,y<i​)∥πT​(⋅∣x,y<i​))].

When extending this approach across model families with different tokenizers, the same text can be assigned different token boundaries and vocabulary entries. Cross-Tokenizer OPD must therefore account for alignment at both the sequence and vocabulary levels. The authors leverage token-group alignment to score and align the decoded student response using both models' tokenizers. The resulting student and teacher response-token sequences are denoted by y=(y1,…,yL)y = (y_1, \dots, y_L)y=(y1​,…,yL​) and v=(v1,…,vn)v = (v_1, \dots, v_n)v=(v1​,…,vn​). By retaining the token offsets shared by both sequences, the response is partitioned into pairs of aligned token groups:

S(y)={(Srθ,SrT)}r=1R.\mathcal{S}(y) = \left\{ (S_r^{\theta}, S_r^{\mathrm{T}}) \right\}_{r=1}^{R}.S(y)={(Srθ​,SrT​)}r=1R​.

These groups are divided into strict 1:1 groups, where a single student token and a single teacher token span the same interval, and mismatch groups, which require multiple tokens on at least one side to construct the same span.

To align the predictions, the authors identify a shared vocabulary V∩=Vθ∩VT\mathcal{V}_{\cap} = \mathcal{V}_{\theta} \cap \mathcal{V}_{\mathrm{T}}V∩​=Vθ​∩VT​ by matching underlying tokens. For strict 1:1 groups, the student and teacher distributions are restricted and renormalized over this shared vocabulary. The strict Cross-Tokenizer objective then sums the reverse KL divergence over strictly aligned positions:

L1:1(θ)=Ex∼D,y∼πθ(⋅∣x)[∑r∈A1:1(y)KL(πˉθ(⋅∣x,y<ir)∥πˉT(⋅∣x,v<jr))].\mathcal{L}_{1:1}(\theta) = \mathbb{E}_{x \sim \mathcal{D}, y \sim \pi_{\theta}(\cdot | x)} \left[ \sum_{r \in \mathcal{A}_{1:1}(y)} \mathrm{KL} \left( \bar{\pi}_{\theta}(\cdot \mid x, y_{<i_r}) \| \bar{\pi}_{\mathrm{T}}(\cdot \mid x, v_{<j_r}) \right) \right].L1:1​(θ)=Ex∼D,y∼πθ​(⋅∣x)​​r∈A1:1​(y)∑​KL(πˉθ​(⋅∣x,y<ir​​)∥πˉT​(⋅∣x,v<jr​​))​.

As shown in the figure below, diagnostic evaluations of this cross-tokenizer approach reveal that strict alignment remains high during training, shared vocabulary captures nearly all probability mass in both models, and excluding mismatch groups yields the best performance.

To evaluate the learning value of the remaining mismatch groups, the authors introduce span supervision. For each mismatch group, the probabilities of the observed token paths are computed for both the student and the teacher. These probabilities are matched using a mean squared error loss over the mismatch groups in a log-probability formulation:

Lspan(θ)=Ex∼D,y∼πθ(⋅∣x)[∑r∈Amis(y)(log⁡qθ(r)−log⁡qT(r))2].\mathcal{L}_{\mathrm{span}}(\theta) = \mathbb{E}_{x \sim \mathcal{D}, y \sim \pi_{\theta}(\cdot | x)} \left[ \sum_{r \in \mathcal{A}_{\mathrm{mis}}(y)} \left( \log q_{\theta}^{(r)} - \log q_{\mathrm{T}}^{(r)} \right)^2 \right].Lspan​(θ)=Ex∼D,y∼πθ​(⋅∣x)​​r∈Amis​(y)∑​(logqθ(r)​−logqT(r)​)2​.

The total training loss combines the strict objective and the span supervision with a weighting factor λ\lambdaλ:

Lλ(θ)=L1:1(θ)+λLspan(θ).\mathcal{L}_{\lambda}(\theta) = \mathcal{L}_{1:1}(\theta) + \lambda \mathcal{L}_{\mathrm{span}}(\theta).Lλ​(θ)=L1:1​(θ)+λLspan​(θ).

As shown in the figure below, the impact of mismatch weights on full-average accuracy across different teacher-student pairs demonstrates that strict supervision alone performs best.

By varying λ\lambdaλ, the authors control the influence of the span loss while maintaining complete supervision coverage. However, empirical results indicate that adding span supervision for mismatch groups tends to reduce downstream accuracy across tested weights, confirming that strict matching retains the most useful supervision.

Experiment

The experiments study cross-tokenizer knowledge distillation on three teacher-student pairs using a shared pool of mathematics and code prompts, with downstream evaluation on math and code benchmarks. They validate that strict token alignment covers most positions on student-generated trajectories despite large static vocabulary mismatch, while adding span supervision for remaining mismatch groups consistently hurts accuracy. The shared vocabulary retains nearly all predictive probability mass, and distillation on a small student-selected top-k subset keeps most of the strict-supervision gains while outperforming four baselines. Gradient diagnostics show that span-loss gradients align poorly with strict-loss gradients and grow in relative magnitude during training, helping explain the negative effect of complete coverage.

Strict token coverage on student trajectories stayed high across the evaluated model pairs even when static vocabulary overlap differed substantially. The pair with the lowest vocabulary Jaccard overlap also had the highest strict coverage under both student and teacher tokenizations. Coverage remained stable across training windows, indicating that large static vocabulary mismatch can coexist with strict alignment at most positions. Strict student and teacher token coverage stayed high across all model pairs despite static vocabulary overlap ranging from about 39% to about 65%. Granite to Phi showed the lowest static vocabulary Jaccard overlap but the highest strict coverage under both tokenizations. Within each pair, strict coverage varied little across training windows, changing by at most 1.43 percentage points for student tokens and 3.75 for teacher tokens.

Strict full-vocabulary distillation consistently improves math and full-average accuracy over baseline alternatives and the base model across all three cross-tokenizer pairs, with the largest full-average gain for Granite-to-Qwen. Code accuracy is more mixed, but strict full remains at or near the best. A compact student-selected subset preserves nearly all of the strict full learning benefit. Strict full supervision leads over all baseline alternatives in math and full-average accuracy across every teacher-student pair. Student-selected top-k subsets retain almost all of the strict full distillation gain, with no consistent further benefit from a larger subset.

The evaluation examines cross-tokenizer distillation across several teacher-student pairs, measuring token coverage and downstream accuracy. Strict token coverage remains high across pairs even when static vocabulary overlap differs substantially, indicating that vocabulary mismatch does not prevent close sequence alignment. Strict full-vocabulary distillation consistently improves math and full-average accuracy over baselines, with the largest full-average gain for Granite-to-Qwen, while code accuracy is more mixed. A compact student-selected subset preserves nearly all of this strict distillation benefit, with no consistent advantage from using a larger subset.


بناء الذكاء الاصطناعي بالذكاء الاصطناعي

من الفكرة إلى الإطلاق — سرّع تطوير الذكاء الاصطناعي الخاص بك مع المساعدة البرمجية المجانية بالذكاء الاصطناعي، وبيئة جاهزة للاستخدام، وأفضل أسعار لوحدات معالجة الرسومات.

البرمجة التعاونية باستخدام الذكاء الاصطناعي
وحدات GPU جاهزة للعمل
أفضل الأسعار

HyperAI Newsletters

اشترك في آخر تحديثاتنا
سنرسل لك أحدث التحديثات الأسبوعية إلى بريدك الإلكتروني في الساعة التاسعة من صباح كل يوم اثنين
مدعوم بواسطة MailChimp