Command Palette
Search for a command to run...
Die On-Policy-Parameter-Update-Richtung liegt der Generalisierung im Post-Training von LLMs zugrunde
Die On-Policy-Parameter-Update-Richtung liegt der Generalisierung im Post-Training von LLMs zugrunde
Shufan Shen Zhongni Hou Junshu Sun Yufei Zhang Wei Lin Guojun Yin Qingming Huang Shuhui Wang
Zusammenfassung
Die starke Generalisierungsleistung von On-Policy-Post-Training-Paradigmen hat Untersuchungen ihres Parameter-Update-Verhaltens motiviert. Diese Studien behandeln die beobachteten Verhaltensweisen jedoch lediglich als Nebenprodukte des On-Policy-Trainings und übersehen ihr Potenzial, als Optimierungsprinzipien zur Verbesserung der Generalisierung anderer Paradigmen wie des überwachten Fine-Tunings (Supervised Fine-Tuning, SFT) zu dienen. Um diese Einschränkung zu beheben, untersuchen wir, ob ein spezifisches On-Policy-Update-Verhalten existiert, das solche Verbesserungen erreichen kann. Zunächst zeigen unsere theoretischen und experimentellen Analysen, dass SFT Parameter entlang konsistenter Richtungen aktualisiert, während das On-Policy-Paradigma die Richtung während des Trainings kontinuierlich anpasst. Dieser Unterschied veranlasst uns, die kumulative Update-Richtung jedes Parameters als vielversprechendes Verhalten in den Fokus zu rücken. Anschließend bewerten wir deren Wirksamkeit zur Verbesserung der Generalisierung, indem wir On-Policy direction-constrained Supervised Fine-Tuning (OPSFT) vorschlagen, das SFT-Updates auf die durch On-Policy-Paradigmen identifizierte Richtung beschränkt. Die starke Leistung von OPSFT zeigt, dass der Generalisierungsvorteil von On-Policy-Paradigmen über die Parameter-Update-Richtung auf SFT übertragen werden kann. Sobald eine solche Richtung identifiziert ist, kann selbst SFT generalisieren, wenn seine Updates auf diese Richtung beschränkt werden. Dieser Befund bietet zwei praktische Vorteile, indem er die starke Generalisierung von On-Policy-Paradigmen mit den Vorteilen von SFT kombiniert, darunter die hohe Trainingseffizienz und die Fähigkeit, hochwertige Trajektorien zu nutzen. Hinsichtlich der Effizienz identifizieren wir Update-Richtungen, die eine starke Generalisierung unterstützen, mit wenigen On-Policy-Trainingsschritten und wenden anschließend OPSFT an, um eine hohe Trainingseffizienz zu erreichen. Zur Nutzung hochwertiger Trajektorien kann OPSFT diese Trajektorien verwenden, um ein post-trainiertes Modell entlang seiner Update-Richtung weiter zu verbessern, ohne die durch das On-Policy-Training erlernte Fähigkeit zu beeinträchtigen.
One-sentence Summary
Researchers at the Chinese Academy of Sciences, the University of Chinese Academy of Sciences, and Meituan propose On-Policy direction-constrained Supervised Fine-Tuning (OPSFT), which constrains supervised fine-tuning (SFT) updates to the cumulative parameter update directions identified by on-policy paradigms, transferring on-policy generalization to efficient SFT while enabling the use of high-quality trajectories.
Key Contributions
- The paper identifies a distinguishing parameter-update behavior: supervised fine-tuning updates parameters along consistent directions, whereas on-policy post-training continuously adjusts update directions during training.
- The paper proposes On-Policy direction-constrained Supervised Fine-Tuning (OPSFT), which constrains SFT updates to the cumulative direction identified by on-policy training; OPSFT achieves performance comparable to the corresponding on-policy paradigm, indicating that the on-policy generalization advantage transfers through update direction.
- The paper demonstrates two practical benefits: a few on-policy training steps identify directions that support strong generalization for efficient OPSFT training, and OPSFT uses newly acquired high-quality trajectories to improve a post-trained model without disrupting previously learned on-policy behavior.
Introduction
On-policy post-training has become an important approach for improving large language model reasoning because it optimizes models on their own generated responses rather than fixed supervised fine-tuning (SFT) targets, and it often generalizes more strongly across tasks. Prior work has examined isolated components such as reverse KL divergence, negative samples, and sparse update locations, but these studies generally treat on-policy update behavior as a byproduct instead of testing whether it can become an optimization principle. That leaves a gap between explaining on-policy generalization and translating those explanations into practical improvements for SFT. The authors bridge this gap by analyzing parameter update direction and find that on-policy training keeps adjusting update directions while SFT stays directionally consistent. They propose On-Policy direction-constrained SFT (OPSFT), which retains SFT but restricts each parameter update to the direction identified by on-policy training. Experiments show that OPSFT substantially outperforms vanilla SFT and reaches performance comparable to the on-policy paradigm, indicating that on-policy update direction can transfer generalization benefits to more efficient SFT.
Method
The authors explore whether an on-policy parameter update behavior can transfer the generalization advantage of on-policy paradigms to Supervised Fine-Tuning (SFT). They investigate the direction, specifically the sign, of parameter updates and its evolution throughout training.
To understand the different behaviors in parameter update direction, the authors compare the gradient formulations of on-policy paradigms and SFT. Given an input prompt x, a response trajectory τ, and the current policy πθ with parameters θ∈Rd, the gradient of the trajectory log-probability with respect to parameters θ is computed as:
sθ(x,τ)=∇θlogπθ(τ∣x)For every input prompt x, sθ(x,τ) has zero conditional expectation:
Eτ∼πθ(⋅∣x)[sθ(x,τ)]=τ∑πθ(τ∣x)∇θlogπθ(τ∣x)=∇θτ∑πθ(τ∣x)=0dwhere 0d∈Rd denotes the zero vector. The policy gradient of on-policy paradigms is represented as:
gon(θ)=Ex,τ∼πθ(⋅∣x)[Aθ(x,τ)sθ(x,τ)]where Aθ(x,τ)∈R is the advantage of trajectory τ for prompt x. The projection of gon(θ) onto an arbitrary sign vector v∈{−1,0,1}d can be formulated as the projection onto v of the covariance between Aθ(x,τ) and sθ(x,τ):
v⊤gon(θ)=Ex[v⊤Covτ∼πθ(⋅∣x)[Aθ(x,τ),sθ(x,τ)]]For SFT, the projection of their gradients along v is:
v⊤gsft(θ)=−Ex[v⊤Eτ∼πteacher(⋅∣x)[sθ(x,τ)]]By comparing these equations, the authors find that SFT updates parameters along the direction of sθ(x,τ) where trajectories are typically sampled from a fixed distribution πteacher. In contrast, on-policy paradigms update parameters toward the direction of the covariance matrix between the advantages and the trajectory gradients. Since trajectories are sampled from the current policy πθ that evolves during training, the distribution of advantages changes with πθ. Consequently, the parameter update direction of on-policy paradigms changes with the evolving advantage distribution throughout training.
To verify this theoretical analysis, the authors measure the cosine similarity among update directions at different training steps. As shown in the figure below:
For interval updates, SFT optimizes towards positively correlated directions, while on-policy paradigms explore nearly orthogonal directions at different training stages. The accumulation of these interval updates leads to substantially different cumulative update directions. For cumulative updates, SFT exhibits highly similar directions throughout training with cosine similarity close to 1.0, whereas on-policy paradigms exhibit positive yet substantially lower correlation with cosine similarity around 0.5. These analyses suggest that unlike SFT, which updates parameters along nearly consistent directions, on-policy paradigms tend to continuously adjust their cumulative parameter update directions throughout training.
Inspired by the continuous efforts to adjust update directions in on-policy paradigms, the authors investigate whether the resulting cumulative on-policy update direction can serve as an effective optimization principle for transferring their generalization advantage to SFT. Specifically, they constrain parameter updates in SFT to the directions identified by on-policy paradigms. Given the parameters before and after on-policy training (θbase,θon), they obtain the sign vector v=sign(θon−θbase)∈{−1,0,1}d that determines the direction of the cumulative update. They constrain the gradient g∈Rd of SFT according to v as follows:
gs=I(sign(−g)=v)⊙gwhere ⊙ denotes the Hadamard product, sign(⋅) is the element-wise sign operator, and I(⋅) represents the indicator function that retains gradient elements whose signs match v while discarding others. The constrained gradient gs is then passed to the optimizer for gradient descent. By constraining the gradient at each training step, the parameter updates consistently follow the directions identified by the on-policy paradigm. In practice, considering that the regularization terms inherent to the optimizer may affect the imposed constraint, the authors further constrain the update direction after each optimizer step. They refer to this On-Policy direction-constrained SFT as OPSFT.
Motivated by the ability of the on-policy update direction to support strong generalization, the authors further leverage this direction to combine the generalization advantage of on-policy paradigms with the advantages of SFT. First, they perform a small number of GRPO steps to identify an update direction that supports strong generalization, and then apply OPSFT to achieve efficient training. Second, given a model post-trained by an on-policy paradigm, they conduct OPSFT along its original update direction, thereby leveraging newly acquired high-quality trajectories to further improve the model without disrupting the capabilities learned during on-policy training.
Experiment
The experiments evaluate whether the generalization benefits of on-policy post-training can be transferred to supervised fine-tuning by sharing parameter update directions, using Qwen and DeepSeek models on math and code reasoning datasets. Compared with vanilla SFT and constraints based only on update locations, constraining SFT updates along on-policy directions achieves performance comparable to or better than GRPO while reducing training time and transferring reasoning capabilities to in-domain and out-of-domain tasks. The identified direction is reusable across datasets within the same domain but not across domains, and it enables further improvement of already post-trained models without the degradation caused by direct SFT. Ablations further show that later-stage directions provide diminishing gains, sparse updates along the constrained direction remain effective, and the resulting reasoning behavior more closely resembles that of on-policy paradigms.
In the out-of-domain evaluations summarized here, OPSFT achieves the highest average performance where it is reported, slightly above GRPO and ahead of vanilla SFT and the base model. Vanilla SFT shows only mixed gains over the base model, and its average can slightly decline at the larger scale, while GRPO improves the average at both shown scales. The paper links these improvements to using the on-policy update direction during SFT. OPSFT leads the Qwen3-4B mean out-of-domain score, with GRPO close behind and vanilla SFT lower. Gains are concentrated in benchmarks like IFEval and HaluEval, while results on ARC, Hellaswag, Winogrande, and PIQA are smaller or mixed.
Across DeepMath benchmarks, OPSFT tends to outperform SFT and DFT in mean accuracy while requiring less training time. It also reduces training time compared with GRPO, often by more than half, while achieving better or comparable generalization. These trends hold across different model scales and architectures reported in the experiments. On Qwen3-1.7B, OPSFT achieves a mean accuracy of 15.11 in 2.3h, improving over SFT, DFT, and GRPO while using the least training time. On DeepSeek-R1-Distill-LLaMA-8B, OPSFT reaches 26.56 mean accuracy in 8.1h, substantially outperforming DFT at 19.90 in 14.1h. Compared with GRPO on Qwen3-8B, OPSFT cuts training time from 19.3h to 8.9h and raises mean accuracy from 40.31 to 41.67.
For Qwen3-1.7B and Qwen3-4B post-trained by GRPO, OPSFT improves mean benchmark accuracy over both the post-trained baseline and vanilla SFT on math reasoning tasks. Vanilla SFT does not reliably preserve post-trained capabilities and can reduce accuracy, as seen with Qwen3-4B, while OPSFT further improves the same model. Results on code tasks are reported as consistent with this trend. OPSFT achieves higher mean accuracy than vanilla SFT across the reported AIME and HMMT reasoning benchmarks. Vanilla SFT can fall below the post-trained baseline, whereas OPSFT improves over it, with the pattern extending to code tasks.
On code tasks with Qwen3-4B and the Eurus dataset, OPSFT achieves the highest mean accuracy while also having the shortest training time among GRPO, SFT, and OPSFT. It improves over SFT on all reported code benchmarks and slightly exceeds GRPO on average despite training in less than half the time. OPSFT reaches the best mean score and the lowest training time across the compared methods. Compared with SFT, OPSFT delivers consistent accuracy gains on HumanEval+, MBPP+, and LCBv6 with slightly shorter training time.
This ablation examines how BF16 and FP32 parameter precision affect update sparsity and accuracy for SFT and OPSFT. OPSFT produces much sparser updates than vanilla SFT under BF16, yet it still improves mean accuracy over vanilla SFT. With FP32 precision, updates become denser and OPSFT achieves the highest mean accuracy, slightly surpassing GRPO while updating a comparable fraction of parameters. Under BF16, OPSFT updates 0.408% of parameters compared with 2.702% for vanilla SFT, but its mean accuracy is higher and substantially above the base model. Under FP32, vanilla SFT becomes dense with about 90% of parameters updated, while OPSFT updates about 9.45%, similar to GRPO's update proportion. FP32 OPSFT attains the highest mean accuracy and edges out GRPO, suggesting that the constrained update direction remains effective at higher precision.
Across out-of-domain, math reasoning, and code evaluations, OPSFT consistently improves over vanilla SFT and the base model while matching or slightly exceeding GRPO, often with substantially shorter training time and sparser updates. Gains are strongest on benchmarks such as IFEval, HaluEval, and math or code tasks, while effects on standard knowledge benchmarks are smaller or mixed. Precision ablations show OPSFT remains effective under BF16 and FP32, updating far fewer parameters than vanilla SFT while achieving the highest mean accuracy and a comparable update density to GRPO under FP32.