HyperAIHyperAI

Command Palette

Search for a command to run...

모든 동전에는 양면이 있다: 대규모 언어 모델의 온폴리시 증류에서 일반화의 이중적 본질에 관하여

초록

온폴리시 증류(OPD)는 학생 모델 자신의 정책에서 샘플링된 궤적을 지도함으로써 교사 모델의 능력을 전이하지만, 대부분의 연구가 OPD를 단일 도메인과 훈련 데이터에 가까운 벤치마크에서만 평가해 왔기 때문에 그 일반화 행동은 여전히 제대로 이해되지 않고 있다. 우리는 도메인 내 분포 이동부터 교차 도메인 전이 및 다중 교사 설정에 이르기까지 한 번에 하나의 일반화 요인만 변화시키는 통제된 연구를 제시한다. 우리는 OPD가 특정 문제에 대한 교사 모델의 답변이 아니라 교사 모델의 추론 행동을 전이한다는 것을 발견했다. 훈련 난이도는 거의 중요하지 않으며, 심지어 교사 모델이 전혀 풀지 못하는 문제도 유용하다. 전이는 교사와 학생 간의 출처 관계에 크게 의존한다. 동일 출처 쌍은 언어, 추론 범위, 심지어 다른 도메인에 걸쳐 학생을 교사에 가깝게 만드는 반면, 교차 출처 쌍은 대체로 훈련된 분포에만 적합한다. 이러한 광범위한 도달 범위는 양날의 검이다. 프롬프트를 도메인 전문가에게 라우팅하는 방식으로는 각 교사의 영향을 제한할 수 없기 때문에, 이들을 결합하면 혼합 의존적인 시소 현상이 그들의 능력 사이에서 발생한다. 이러한 결과는 OPD가 언제 일반화되는지를 명확히 하고 다중 교사 OPD를 진단하는 데 유용한 관점을 제공한다.

One-sentence Summary

Researchers from the University of Science and Technology of China, Peking University, and other institutions demonstrate that on-policy distillation transfers a teacher's reasoning behavior rather than specific answers, with same-origin teacher-student pairs enabling broad cross-domain generalization while cross-origin pairs mainly fit the training distribution, and that multi-teacher distillation produces a mixture-dependent seesaw effect among capabilities, thereby clarifying when OPD generalizes.

Key Contributions

  • A controlled study systematically varies generalization factors in on-policy distillation (OPD) across math, code, science, and instruction-following domains, isolating in-domain shifts in problem difficulty, language, and reasoning horizon, cross-domain transfer, and multi-teacher settings.
  • OPD transfers a teacher’s reasoning behavior rather than its answers to specific problems: training problem difficulty barely matters, and even problems the teacher never solves are useful.
  • Generalization depends strongly on the teacher-student origin relationship: same-origin pairs achieve broad transfer across languages, reasoning horizons, and even other domains, while cross-origin pairs mostly fit the training distribution; in multi-teacher OPD, routing cannot confine each teacher’s influence, yielding a mixture-dependent seesaw among their capabilities.

Introduction

The authors investigate on-policy distillation (OPD), a technique that transfers capabilities from strong teacher models to smaller student models by supervising the student on its own generated trajectories. While prior work demonstrates OPD improves performance on target tasks, it fails to distinguish whether the student merely fits the training distribution or acquires broader reasoning patterns that generalize to new settings. This gap is critical for understanding how OPD works and for designing effective multi-teacher integration strategies. The authors conduct a controlled study that systematically varies one generalization factor at a time, revealing that OPD transfers the teacher’s reasoning behavior rather than specific problem solutions, that same-origin teacher-student pairs transfer broadly across languages, horizons, and even domains, and that this broad transfer becomes a double-edged sword in multi-teacher OPD, where prompt routing cannot isolate each teacher’s influence, leading to a mixture-dependent seesaw among capabilities.

Experiment

This work evaluates on-policy distillation with single and multiple teachers across math, code, science, and instruction following. It finds that training-problem difficulty has little effect, while dynamically discarding problems the student already solves gives small consistent gains. Same-origin teacher-student pairs transfer capabilities across language, reasoning horizon, and domains much more effectively than cross-origin pairs, which fit mainly the trained distribution. Because a teacher's influence reaches beyond its assigned domain, multi-teacher routing cannot isolate domain experts; changing teacher mixture ratios creates a seesaw effect, and same-origin teachers exert a stronger pull by aligning the student policy as a whole.

Dynamically discarding only the problems the student already solves (pass-rate in [0,1)) yields a small but consistent average improvement across six math benchmarks, while filtering to only unsolved or only solved problems does not help. This suggests that removing already-mastered problems prevents redundant realignment without harming generalization. For Polaris-7B → DS-distill-1.5B, discarding fully solved problems raised average accuracy from 41.4% to 42.0% (+0.6 pp). Restricting to only unsolved (pass-rate=0) or only solved (pass-rate=1) problems resulted in average accuracy equal to or below the no-filtering baseline. The benefit of discarding solved problems was consistent across both teacher-student pairs, with gains on most individual benchmarks.

Varying the prompt mixture ratio between a math-specialist teacher and a science/IF-specialist teacher in multi-teacher online policy distillation shifts the student's math accuracy toward the teacher that receives more prompts. When the math teacher JustRL-1.5B dominates, average math accuracy rises slightly; when the science/IF teacher Nemotron-1.5B dominates, math accuracy falls. This demonstrates the cross-domain seesaw effect, where a teacher's influence pulls performance in domains beyond its assigned teaching area. Increasing the share of prompts given to the math teacher improves the student's math accuracy, while increasing the science/IF teacher's share degrades it. The highest average math accuracy (27.1%) is achieved when the math teacher receives the largest prompt share (J/N=25/8), and the lowest (25.1%) when the science/IF teacher dominates (J/N=2/25). Relative to an equal mixture, a math-teacher-heavy ratio boosts math accuracy by 0.8 percentage points, whereas a science-teacher-heavy ratio reduces it by 1.2 points.

Filtering out problems the student already solves during online policy distillation yields a small but consistent accuracy improvement by preventing redundant realignment, while restricting to only unsolved or only solved problems does not help. Varying the prompt mixture between a math-specialist teacher and a science/instruction-following teacher demonstrates a cross-domain seesaw effect, where the student's math performance shifts toward the teacher that receives more prompts, even beyond that teacher's assigned domain. These results validate that selective problem filtering and teacher prompt allocation are effective levers for guiding student performance in multi-teacher distillation.


AI로 AI 구축

아이디어에서 출시까지 — 무료 AI 코코딩, 즉시 사용 가능한 환경, 최적의 GPU 가격으로 AI 개발을 가속화하세요.

AI 협업 코딩
바로 사용 가능한 GPU
최적의 가격

HyperAI Newsletters

최신 정보 구독하기
한국 시간 매주 월요일 오전 9시 에 이번 주의 최신 업데이트를 메일로 발송합니다
이메일 서비스 제공: MailChimp