HyperAIHyperAI

Command Palette

Search for a command to run...

Post-Training Leaves Behavioral Shadows on Unrelated Decisions

Ziyang Zhang Yubin Jing Yuanhao Zeng Yuyao Li Haofan Wang Yichen Gong

Abstract

We find that language models can transfer capabilities through task-unrelated text. Post-training typically improves language models using task-specific data. Prior work on subliminal learning shows that information about these updates can pass through unrelated generations, but has largely focused on traits or preferences using extensive teacher outputs. We introduce Active Taskless Distillation (ATD), which achieves capability transfer using only a single word from the teacher per prompt. ATD probes the behavioral shadow of post-training by selecting prompts where the teacher and student’s shared public ancestor is nearly indiferent between two ordinary words. A student initialized from this ancestor learns solely from the resulting prompt–word pairs, without target-task examples, teacher logits, or teacher parameters. In the primary coding experiment with Qwen2.5-1.5B, 5,664 single-word teacher responses yield a 5.34 pp gain on HumanEval+ over an exact nuisance-matched control that disrupts prompt–response pairings. Further experiments show transfer in scientific knowledge, commonsense reasoning, and reading comprehension, with positive mean gains across additional model generations, sizes, and families. Functional analyses show that the learned signal is source-specific and composable, and that its strength tracks the teacher’s update strength. Code is available at: github.com/myboker/ATD.

One-sentence Summary

Researchers at Peking University, Georgia Institute of Technology, and other institutions propose Active Taskless Distillation (ATD), which transfers capabilities through single-word teacher responses by selecting prompts where the shared public ancestor is nearly indifferent between two ordinary words, and in the primary coding experiment with Qwen2.5-1.5B, 5,664 such responses yield a 5.34 pp HumanEval+ gain and transfer across reasoning, knowledge, and comprehension tasks.

Key Contributions

  • The paper introduces Active Taskless Distillation (ATD), a method that transfers capabilities from a post-trained teacher to a same-ancestor student using only one ordinary word per prompt as supervision, selected where the shared public ancestor is nearly indifferent between two words and without target-task examples, teacher logits, or teacher parameters.
  • In the primary coding experiment with Qwen2.5-1.5B, 5,664 single-word teacher responses yield a 5.34 percentage point gain on HumanEval+ over an exact nuisance-matched control; further experiments show transfer in scientific knowledge, commonsense reasoning, and reading comprehension, with positive mean gains across additional model generations, sizes, and families.
  • Functional analyses show that the learned signal is source-specific and composable, that its strength tracks the teacher’s update strength, and that an early teacher-aligned change is measurable in the student before target-task gains reach statistical significance.

Introduction

Post-training is usually assessed through improvements on the target task, but an update can also shift a model’s preferences on unrelated inputs, a phenomenon the authors call the behavioral shadow. This matters because such unintended changes may reveal information about private model adaptations without access to updated parameters or target-task data. Prior work on subliminal learning showed that visible content does not fully determine what a training example exposes, but existing distillation and extraction methods typically rely on teacher logits, target-task responses, or longer generative outputs. The authors introduce Active Taskless Distillation (ATD), which uses the public ancestor to find prompts where two ordinary words are nearly tied, observes which word the privately updated teacher selects, and trains a student on these one-bit choices. Their results show that this taskless signal carries capability-relevant information, improving student performance on code generation and multiple-choice benchmarks despite using no target-task examples or teacher probabilities.

Method

The authors introduce Active Taskless Distillation (ATD), a black-box distillation procedure that extracts capability-relevant signal from a private teacher without using target-task data. The setting assumes a public ancestor model M0M_0M0​ with parameters θ0\theta_0θ0​ and a private teacher MTM_TMT​ obtained by post-training M0M_0M0​ on a target task:

θT=θ0+δ,\theta_T = \theta_0 + \delta,θT​=θ0​+δ,

where δ\deltaδ is the unknown post-training update. The student is initialized from the same θ0\theta_0θ0​. During distillation, the teacher is available only through a black-box interface that returns one next token per query by greedy decoding over the full vocabulary, and the student receives no target-task data.

ATD is built around the notion of carriers and the behavioral shadow. A carrier is a prompt-response pair whose visible text is unrelated to the target capability. On such task-unrelated inputs, the private teacher may still behave differently from its public ancestor. These differences are the behavioral shadow of the private update. To determine whether the shadow carries useful signal, a same-ancestor student is trained on the resulting prompt-word pairs and evaluated on the target task. Improvement over matched controls that disrupt the prompt-response correspondence indicates capability transfer through these observations.

The method uses near-tie prompts because they can expose the private update through small preference changes. For a carrier prompt xix_ixi​ with two ordinary single-token candidates aia_iai​ and bib_ibi​, the logit difference is defined as

mi(θ)=zθ(ai∣xi)−zθ(bi∣xi),m_i(\theta) = z_\theta(a_i \mid x_i) - z_\theta(b_i \mid x_i),mi​(θ)=zθ​(ai​∣xi​)−zθ​(bi​∣xi​),

where zθ(w∣xi)z_\theta(w \mid x_i)zθ​(w∣xi​) is the next-token logit for word www. A first-order expansion around θ0\theta_0θ0​ gives

mi(θ0+δ)≈mi(θ0)+gi⊤δ,gi=∇θmi(θ0).m_i(\theta_0 + \delta) \approx m_i(\theta_0) + g_i^\top \delta, \qquad g_i = \nabla_\theta m_i(\theta_0).mi​(θ0​+δ)≈mi​(θ0​)+gi⊤​δ,gi​=∇θ​mi​(θ0​).

Near a tie, ∣mi(θ0)∣|m_i(\theta_0)|∣mi​(θ0​)∣ is small, so a small change in relative preference can reverse the selected word. The observed teacher choice is therefore modeled as

yi≈sign(mi(θ0)+gi⊤δ).y_i \approx \mathrm{sign}\big(m_i(\theta_0) + g_i^\top \delta\big).yi​≈sign(mi​(θ0​)+gi⊤​δ).

Each retained teacher response provides a one-bit observation of the behavioral shadow: which word the teacher prefers, but not the magnitude of the preference change.

ATD proceeds in three main stages. First, the authors construct candidate prompts that ask the model to choose between two ordinary single-token words aia_iai​ and bib_ibi​. Using only the public ancestor M0M_0M0​, they compute the normalized probability of bib_ibi​ over the two candidates as

qi=exp⁡z0(bi∣xi)exp⁡z0(ai∣xi)+exp⁡z0(bi∣xi).q_i = \frac{\exp z_0(b_i \mid x_i)} {\exp z_0(a_i \mid x_i) + \exp z_0(b_i \mid x_i)}.qi​=expz0​(ai​∣xi​)+expz0​(bi​∣xi​)expz0​(bi​∣xi​)​.

Prompts are retained when

∣qi−0.5∣≤ϵ,|q_i - 0.5| \le \epsilon,∣qi​−0.5∣≤ϵ,

with ϵ=0.02\epsilon = 0.02ϵ=0.02. The authors also require that the highest-probability token under the public model over the full vocabulary is one of the two candidates. This ensures the near-tie involves the model's actual output choice.

Second, the selected prompts and candidate pairs are fixed before observing any teacher responses. Each prompt is queried once, and the teacher returns its highest-probability next token over the full vocabulary. The response wiTw_i^TwiT​ is retained only if it belongs to {ai,bi}\{a_i, b_i\}{ai​,bi​}; otherwise the example is discarded without resampling. Each retained binary choice forms a training pair (xi,wiT)(x_i, w_i^T)(xi​,wiT​), even when neither the prompt nor the response mentions the target task.

Third, the student is initialized from the same public ancestor M0M_0M0​. For the nnn retained prompt-word pairs, the authors minimize full-vocabulary cross-entropy:

LCE(θ)=−1n∑i=1nlog⁡pθ(wiT∣xi).\mathcal{L}_{\mathrm{CE}}(\theta) = -\frac{1}{n} \sum_{i=1}^{n} \log p_\theta(w_i^T \mid x_i).LCE​(θ)=−n1​i=1∑n​logpθ​(wiT​∣xi​).

Only the response token contributes to the loss, while the prompt supplies context.

The authors also explore two refinements of the same instrument. The first is a reference-relative pair-margin (RPM) objective, which learns the residual preference introduced by δ\deltaδ rather than the preference already present in M0M_0M0​. Let wˉi\bar{w}_iwˉi​ denote the other word in the pair. RPM subtracts the frozen public margin between the two candidate words:

LRPM=−Eilog⁡σ(β[log⁡pθ(wiT∣xi)pθ(wˉi∣xi)−log⁡p0(wiT∣xi)p0(wˉi∣xi)]).\mathcal{L}_{\mathrm{RPM}} = -\mathbb{E}_i \log \sigma \left( \beta \left[ \log \frac{p_\theta(w_i^T \mid x_i)} {p_\theta(\bar{w}_i \mid x_i)} - \log \frac{p_0(w_i^T \mid x_i)} {p_0(\bar{w}_i \mid x_i)} \right] \right).LRPM​=−Ei​logσ(β[logpθ​(wˉi​∣xi​)pθ​(wiT​∣xi​)​−logp0​(wˉi​∣xi​)p0​(wiT​∣xi​)​]).

The second extension widens each observation from a single hard token to a KKK-word answer yi=(yi,1,…,yi,K)y_i = (y_{i,1}, \dotsc, y_{i,K})yi​=(yi,1​,…,yi,K​) on the same carriers. The student is then trained by ordinary full-vocabulary cross-entropy over the KKK positions:

Lmulti=−1n∑i1K∑k=1Klog⁡pθ(yi,k∣xi,yi,<k),\mathcal{L}_{\mathrm{multi}} = -\frac{1}{n} \sum_{i} \frac{1}{K} \sum_{k=1}^{K} \log p_\theta(y_{i,k} \mid x_i, y_{i,<k}),Lmulti​=−n1​i∑​K1​k=1∑K​logpθ​(yi,k​∣xi​,yi,<k​),

which reduces to the single-token objective when K=1K=1K=1. These extensions are exploratory refinements of the same near-boundary instrument rather than separate methods.

Experiment

The experiments evaluate an implicit distillation setup in which a student initialized from a public ancestor is trained only on single-token teacher responses to target-unrelated, near-tie carrier prompts, then tested on coding, math, and multiple-choice benchmarks against matched controls. The central finding is that such training recovers a private teacher capability on HumanEval+, with the effect surviving exact nuisance-matched and shuffled controls, and it remains robust across fresh acquisitions, re-trained teachers, different adaptation regimes, model families, and several unrelated tasks. Analyses show that transfer comes from actively querying the ancestor's decision boundary rather than query volume, is source-specific and composable, and does not follow from large teacher gains alone, indicating the channel carries expressible capability rather than raw teacher advantage; multi-token observations extend the same effect with task-dependent strength.

On HumanEval+ pass@1 with Qwen2.5-1.5B students, the signal model reaches 51.22% and exceeds every identification control. The gaps over the exact nuisance-matched, shuffled-label, and teacher-free controls are roughly five points, and the primary controls have 95% confidence intervals excluding zero. A single-seed random-marginal diagnostic also trails the signal by about seven points. The signal outperforms the exact nuisance-matched control by about five points, with a 95% confidence interval that excludes zero. Teacher-free and shuffled-label controls also score lower than the signal by roughly five points, so generic carrier fine-tuning and teacher label statistics do not explain the gain.

Signal-control gaps are positive across all adaptation regimes, with confidence intervals that lie above zero. The estimated effect is largest with LoRA r64 and smallest with full SFT, showing robustness across adaptation choices. Every adaptation regime tested shows a positive signal-control effect with confidence intervals entirely above zero. Larger LoRA ranks correspond to larger signal-control gaps, while full SFT yields the smallest gap among the regimes.

Across seven target tasks spanning code generation, scientific knowledge, commonsense reasoning, and reading comprehension, students trained on task-unrelated teacher responses outperformed matched controls in every setting. All transfer gains had confidence intervals excluding zero, ranging from under one percentage point to about five percentage points. The largest transfer appeared on code generation, while several tasks with larger teacher advantages showed smaller or mid-range improvements. Every evaluated target task showed a positive transfer effect over the teacher-label shuffle control, with gains ranging from under one point to about five points. A larger teacher advantage did not necessarily translate into larger transfer; HellaSwag and ScienceQA had some of the largest teacher gains but only modest transfer improvements compared with code generation.

Under a matched budget, active acquisition outperforms passive acquisition for both teacher queries and training rows. The reported differences are positive with confidence intervals above zero, indicating a consistent advantage. The gain is slightly larger for training rows than for teacher queries. Active acquisition yields a positive difference over passive acquisition for teacher queries. Active acquisition yields a positive difference over passive acquisition for training rows, with a somewhat larger gain than for teacher queries.

Analyses of non-elicitable teacher advantages show that a large teacher gain does not by itself transfer to the student. A teacher that memorized HumanEval+ solutions produced a large target-task gap but the student signal remained below its base, and a substitution cipher teacher produced no improvement over a zero base. These cases suggest transfer depends on capabilities the student can be elicited to express rather than the teacher’s raw advantage. A memorized HumanEval+ teacher achieved a large gain on its endpoint, but the ATD student did not improve over its base model. A substitution cipher teacher reached a near-perfect gap, yet both base and student remained at zero, indicating no transfer of that non-elicitable capability.

The experiments evaluate a signal-based approach using Qwen2.5-1.5B students across code generation, scientific knowledge, commonsense reasoning, and reading comprehension, comparing it with nuisance-matched, shuffled-label, teacher-free, and passive acquisition controls. The signal model consistently outperforms these controls, with positive gaps across adaptation regimes and target tasks, and active acquisition improves over passive acquisition for both teacher queries and training rows. Gains are not explained by generic carrier fine-tuning or teacher label statistics, while non-elicitable teacher advantages such as memorized solutions or cipher skills do not transfer, indicating that transfer depends on capabilities the student can actually express.


Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp