Command Palette
Search for a command to run...
Le post-entraînement laisse des ombres comportementales sur des décisions sans rapport
Le post-entraînement laisse des ombres comportementales sur des décisions sans rapport
Ziyang Zhang Yubin Jing Yuanhao Zeng Yuyao Li Haofan Wang Yichen Gong
Résumé
Nous constatons que les modèles de langage peuvent transférer des capacités à travers du texte sans lien avec la tâche. Le post-entraînement améliore généralement les modèles de langage à l'aide de données spécifiques à la tâche. Des travaux antérieurs sur l'apprentissage subliminal montrent que des informations sur ces mises à jour peuvent transiter par des générations non liées, mais se sont largement concentrés sur des traits ou des préférences en utilisant de nombreuses sorties de l'enseignant. Nous introduisons la distillation active sans tâche (Active Taskless Distillation, ATD), qui réalise un transfert de capacités en utilisant un seul mot de l'enseignant par invite. ATD sonde l'ombre comportementale du post-entraînement en sélectionnant des invites où l'ancêtre public commun à l'enseignant et à l'élève est presque indifférent entre deux mots ordinaires. Un élève initialisé à partir de cet ancêtre apprend uniquement à partir des paires invite–mot ainsi obtenues, sans exemples de la tâche cible, sans logits ni paramètres de l'enseignant. Dans l'expérience principale de codage avec Qwen2.5-1.5B, 5 664 réponses d'un seul mot de l'enseignant produisent un gain de 5,34 points de pourcentage sur HumanEval+ par rapport à un contrôle exact apparié sur les facteurs de nuisance qui rompt les appariements invite–réponse. D'autres expériences montrent un transfert en connaissances scientifiques, en raisonnement de bon sens et en compréhension de lecture, avec des gains moyens positifs à travers d'autres générations de modèles, tailles et familles. Des analyses fonctionnelles montrent que le signal appris est spécifique à la source et composable, et que son intensité suit l'intensité de la mise à jour de l'enseignant. Le code est disponible à l'adresse : github.com/myboker/ATD.
One-sentence Summary
Researchers at Peking University, Georgia Institute of Technology, and other institutions propose Active Taskless Distillation (ATD), which transfers capabilities through single-word teacher responses by selecting prompts where the shared public ancestor is nearly indifferent between two ordinary words, and in the primary coding experiment with Qwen2.5-1.5B, 5,664 such responses yield a 5.34 pp HumanEval+ gain and transfer across reasoning, knowledge, and comprehension tasks.
Key Contributions
- The paper introduces Active Taskless Distillation (ATD), a method that transfers capabilities from a post-trained teacher to a same-ancestor student using only one ordinary word per prompt as supervision, selected where the shared public ancestor is nearly indifferent between two words and without target-task examples, teacher logits, or teacher parameters.
- In the primary coding experiment with Qwen2.5-1.5B, 5,664 single-word teacher responses yield a 5.34 percentage point gain on HumanEval+ over an exact nuisance-matched control; further experiments show transfer in scientific knowledge, commonsense reasoning, and reading comprehension, with positive mean gains across additional model generations, sizes, and families.
- Functional analyses show that the learned signal is source-specific and composable, that its strength tracks the teacher’s update strength, and that an early teacher-aligned change is measurable in the student before target-task gains reach statistical significance.
Introduction
Post-training is usually assessed through improvements on the target task, but an update can also shift a model’s preferences on unrelated inputs, a phenomenon the authors call the behavioral shadow. This matters because such unintended changes may reveal information about private model adaptations without access to updated parameters or target-task data. Prior work on subliminal learning showed that visible content does not fully determine what a training example exposes, but existing distillation and extraction methods typically rely on teacher logits, target-task responses, or longer generative outputs. The authors introduce Active Taskless Distillation (ATD), which uses the public ancestor to find prompts where two ordinary words are nearly tied, observes which word the privately updated teacher selects, and trains a student on these one-bit choices. Their results show that this taskless signal carries capability-relevant information, improving student performance on code generation and multiple-choice benchmarks despite using no target-task examples or teacher probabilities.
Method
The authors introduce Active Taskless Distillation (ATD), a black-box distillation procedure that extracts capability-relevant signal from a private teacher without using target-task data. The setting assumes a public ancestor model M0 with parameters θ0 and a private teacher MT obtained by post-training M0 on a target task:
θT=θ0+δ,where δ is the unknown post-training update. The student is initialized from the same θ0. During distillation, the teacher is available only through a black-box interface that returns one next token per query by greedy decoding over the full vocabulary, and the student receives no target-task data.
ATD is built around the notion of carriers and the behavioral shadow. A carrier is a prompt-response pair whose visible text is unrelated to the target capability. On such task-unrelated inputs, the private teacher may still behave differently from its public ancestor. These differences are the behavioral shadow of the private update. To determine whether the shadow carries useful signal, a same-ancestor student is trained on the resulting prompt-word pairs and evaluated on the target task. Improvement over matched controls that disrupt the prompt-response correspondence indicates capability transfer through these observations.
The method uses near-tie prompts because they can expose the private update through small preference changes. For a carrier prompt xi with two ordinary single-token candidates ai and bi, the logit difference is defined as
mi(θ)=zθ(ai∣xi)−zθ(bi∣xi),where zθ(w∣xi) is the next-token logit for word w. A first-order expansion around θ0 gives
mi(θ0+δ)≈mi(θ0)+gi⊤δ,gi=∇θmi(θ0).Near a tie, ∣mi(θ0)∣ is small, so a small change in relative preference can reverse the selected word. The observed teacher choice is therefore modeled as
yi≈sign(mi(θ0)+gi⊤δ).Each retained teacher response provides a one-bit observation of the behavioral shadow: which word the teacher prefers, but not the magnitude of the preference change.
ATD proceeds in three main stages. First, the authors construct candidate prompts that ask the model to choose between two ordinary single-token words ai and bi. Using only the public ancestor M0, they compute the normalized probability of bi over the two candidates as
qi=expz0(ai∣xi)+expz0(bi∣xi)expz0(bi∣xi).Prompts are retained when
∣qi−0.5∣≤ϵ,with ϵ=0.02. The authors also require that the highest-probability token under the public model over the full vocabulary is one of the two candidates. This ensures the near-tie involves the model's actual output choice.
Second, the selected prompts and candidate pairs are fixed before observing any teacher responses. Each prompt is queried once, and the teacher returns its highest-probability next token over the full vocabulary. The response wiT is retained only if it belongs to {ai,bi}; otherwise the example is discarded without resampling. Each retained binary choice forms a training pair (xi,wiT), even when neither the prompt nor the response mentions the target task.
Third, the student is initialized from the same public ancestor M0. For the n retained prompt-word pairs, the authors minimize full-vocabulary cross-entropy:
LCE(θ)=−n1i=1∑nlogpθ(wiT∣xi).Only the response token contributes to the loss, while the prompt supplies context.
The authors also explore two refinements of the same instrument. The first is a reference-relative pair-margin (RPM) objective, which learns the residual preference introduced by δ rather than the preference already present in M0. Let wˉi denote the other word in the pair. RPM subtracts the frozen public margin between the two candidate words:
LRPM=−Eilogσ(β[logpθ(wˉi∣xi)pθ(wiT∣xi)−logp0(wˉi∣xi)p0(wiT∣xi)]).The second extension widens each observation from a single hard token to a K-word answer yi=(yi,1,…,yi,K) on the same carriers. The student is then trained by ordinary full-vocabulary cross-entropy over the K positions:
Lmulti=−n1i∑K1k=1∑Klogpθ(yi,k∣xi,yi,<k),which reduces to the single-token objective when K=1. These extensions are exploratory refinements of the same near-boundary instrument rather than separate methods.
Experiment
The experiments evaluate an implicit distillation setup in which a student initialized from a public ancestor is trained only on single-token teacher responses to target-unrelated, near-tie carrier prompts, then tested on coding, math, and multiple-choice benchmarks against matched controls. The central finding is that such training recovers a private teacher capability on HumanEval+, with the effect surviving exact nuisance-matched and shuffled controls, and it remains robust across fresh acquisitions, re-trained teachers, different adaptation regimes, model families, and several unrelated tasks. Analyses show that transfer comes from actively querying the ancestor's decision boundary rather than query volume, is source-specific and composable, and does not follow from large teacher gains alone, indicating the channel carries expressible capability rather than raw teacher advantage; multi-token observations extend the same effect with task-dependent strength.
On HumanEval+ pass@1 with Qwen2.5-1.5B students, the signal model reaches 51.22% and exceeds every identification control. The gaps over the exact nuisance-matched, shuffled-label, and teacher-free controls are roughly five points, and the primary controls have 95% confidence intervals excluding zero. A single-seed random-marginal diagnostic also trails the signal by about seven points. The signal outperforms the exact nuisance-matched control by about five points, with a 95% confidence interval that excludes zero. Teacher-free and shuffled-label controls also score lower than the signal by roughly five points, so generic carrier fine-tuning and teacher label statistics do not explain the gain.
Signal-control gaps are positive across all adaptation regimes, with confidence intervals that lie above zero. The estimated effect is largest with LoRA r64 and smallest with full SFT, showing robustness across adaptation choices. Every adaptation regime tested shows a positive signal-control effect with confidence intervals entirely above zero. Larger LoRA ranks correspond to larger signal-control gaps, while full SFT yields the smallest gap among the regimes.
Across seven target tasks spanning code generation, scientific knowledge, commonsense reasoning, and reading comprehension, students trained on task-unrelated teacher responses outperformed matched controls in every setting. All transfer gains had confidence intervals excluding zero, ranging from under one percentage point to about five percentage points. The largest transfer appeared on code generation, while several tasks with larger teacher advantages showed smaller or mid-range improvements. Every evaluated target task showed a positive transfer effect over the teacher-label shuffle control, with gains ranging from under one point to about five points. A larger teacher advantage did not necessarily translate into larger transfer; HellaSwag and ScienceQA had some of the largest teacher gains but only modest transfer improvements compared with code generation.
Under a matched budget, active acquisition outperforms passive acquisition for both teacher queries and training rows. The reported differences are positive with confidence intervals above zero, indicating a consistent advantage. The gain is slightly larger for training rows than for teacher queries. Active acquisition yields a positive difference over passive acquisition for teacher queries. Active acquisition yields a positive difference over passive acquisition for training rows, with a somewhat larger gain than for teacher queries.
Analyses of non-elicitable teacher advantages show that a large teacher gain does not by itself transfer to the student. A teacher that memorized HumanEval+ solutions produced a large target-task gap but the student signal remained below its base, and a substitution cipher teacher produced no improvement over a zero base. These cases suggest transfer depends on capabilities the student can be elicited to express rather than the teacher’s raw advantage. A memorized HumanEval+ teacher achieved a large gain on its endpoint, but the ATD student did not improve over its base model. A substitution cipher teacher reached a near-perfect gap, yet both base and student remained at zero, indicating no transfer of that non-elicitable capability.
The experiments evaluate a signal-based approach using Qwen2.5-1.5B students across code generation, scientific knowledge, commonsense reasoning, and reading comprehension, comparing it with nuisance-matched, shuffled-label, teacher-free, and passive acquisition controls. The signal model consistently outperforms these controls, with positive gaps across adaptation regimes and target tasks, and active acquisition improves over passive acquisition for both teacher queries and training rows. Gains are not explained by generic carrier fine-tuning or teacher label statistics, while non-elicitable teacher advantages such as memorized solutions or cipher skills do not transfer, indicating that transfer depends on capabilities the student can actually express.