HyperAIHyperAI

Command Palette

Search for a command to run...

UNIEVO-VL : une recette d’entraînement par auto-distillation on-policy pour l’auto-amélioration des modèles multimodaux

Résumé

Les modèles multimodaux modernes réunissent la génération et la compréhension dans un système unique, ce qui leur permet de fournir leurs propres retours et d’apprendre de ceux-ci. Motivés par cette capacité unifiée, nous présentons UniEvo-VL, un cadre auto-évolutif permettant aux modèles multimodaux d’apprendre de ce retour d’auto-correction constructif pendant le calcul au moment du test. Au lieu de recourir à un enseignant séparé, souvent plus grand, nous exploitons leurs auto-critiques comme information privilégiée et demandons à un unique modèle multimodal de jouer à la fois le rôle d’enseignant et celui d’élève avec des contextes différents. L’élève ne voit que la question originale, tandis que l’enseignant est conditionné par la critique privilégiée. L’entraînement minimise ensuite la divergence état par état entre leurs distributions de diffusion débruitantes le long des trajectoires d’échantillonnage propres à l’élève. Les expériences montrent qu’UniEvo-VL améliore les capacités de génération d’images des modèles multimodaux, tout en maintenant leur sensibilité à une information de réflexion supplémentaire. Plus précisément, nous nous appuyons sur le modèle open source Qwen-image-2512 et observons un gain de performance significatif, de 0,747 à 0,808 sur GenEval et de 32,97 à 35,53 sur GenEval2 Soft-TIFA. En outre, des tentatives avec des critiques externes plus puissants, par exemple GPT5.6-Luna, montrent que les modèles multimodaux dotés de fortes capacités de jugement peuvent anticiper un plafond d’auto-évolution plus élevé. Enfin, des résultats mitigés en rendu de texte montrent que nos auto-améliorations peuvent ne pas être uniformes selon les tâches. Notre étude vise à éclairer la ligne de recherche actuelle et très active sur l’auto-amélioration récursive afin d’améliorer l’expérience utilisateur lors de l’utilisation de modèles multimodaux sans supervision ni guidage externes.

One-sentence Summary

Stanford University, Johns Hopkins University, and colleagues introduce UniEvo-VL, a self-evolving framework that uses a single multimodal model as both teacher and student by conditioning the teacher on self-critiques and minimizing per-state divergence between denoising diffusion distributions, improving image generation in Qwen-image-2512 from 0.7470.7470.747 to 0.8080.8080.808 on GenEval without external supervision.

Key Contributions

  • UniEvo-VL introduces a critique-conditioned on-policy self-distillation framework in which one multimodal model acts as both teacher and student. The teacher conditions on self-generated corrective critiques while the student sees only the original prompt, and training matches denoising distributions along the student’s own sampling trajectories without corrected-image targets or reward-based policy optimization.
  • The work separates learned improvements from inference-time correction by comparing direct and reflection-assisted generation before and after training. This paired evaluation distinguishes gains retained in the initial model from the additional benefit of reflection at test time.
  • Experiments on Qwen-image-2512 improve GenEval from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53. Further analysis of critic choice and post-revision verification shows that attempts with more powerful external critics such as GPT5.6-Luna suggest a higher self-evolving ceiling, while gains are uneven across text-rendering tasks.

Introduction

Unified multimodal models can both generate and understand images, creating an opportunity for self-evolution through self-critique. In compositional image generation and visual text rendering, prior self-improvement methods rely on selecting self-generated outputs for fine-tuning, training with reflection triplets, or converting multimodal assessments into rewards. However, these methods do not directly transfer corrective feedback into the policy that generates from the original prompt along its own sampling trajectories; a critique states what should be fixed but not how intermediate denoising predictions should change. The authors address this gap with UniEvo-VL, a self-distillation framework in which a teacher model conditions on a critique-revised prompt while the student model observes only the original request, and their predictions are matched on the student’s own trajectories. This converts visual critique into dense state-wise supervision without corrected-image targets or reward-based policy optimization.

Method

UniEvo-VL is built around a single multimodal model Mθ\mathcal{M}_\thetaMθ​ that supports two operating modes: generation and understanding. In generation mode, the model maps a prompt ppp and sampled noise ϵ\epsilonϵ to an image:

I=Mθgen(p,ϵ).I = \mathcal{M}_\theta^{\mathrm{gen}}(p, \epsilon).I=Mθgen​(p,ϵ).

In understanding mode, the model evaluates an image III against a conditioning prompt ppp and returns a binary acceptance decision aaa together with a discrepancy description ccc:

(c,a)=Mθund(p,I).(c, a) = \mathcal{M}_\theta^{\mathrm{und}}(p, I).(c,a)=Mθund​(p,I).

When a=1a=1a=1, the image is considered aligned with the prompt and does not require revision. When a=0a=0a=0, the image is considered flawed, and ccc provides corrective feedback. Empty or unusable critique responses are treated as invalid and removed from the loop.

The authors connect generation and understanding through corrective conditioning. They define two policies. The student policy Mθ\mathcal{M}_\thetaMθ​ observes only the vanilla prompt ppp. The teacher policy is an exponential moving average version Mθˉ\mathcal{M}_{\bar\theta}Mθˉ​ of the same model, and it is allowed to observe privileged correction information. This setup enables the teacher to generate distillation targets that reflect the corrections while the student must learn to achieve the same behavior without access to the privileged context.

For an image that receives a valid rejection with a=0a=0a=0, the understanding mode synthesizes a privileged prompt p~\widetilde{p}p​ from the original prompt ppp and the corrective feedback ccc:

p~=Mθund(p,c).\widetilde{p} = \mathcal{M}_\theta^{\mathrm{und}}(p, c).p​=Mθund​(p,c).

The privileged prompt restates the requested scene while explicitly resolving the discrepancies identified in ccc. For example, if the original prompt requests two red cups but the model generates three cups, p~\widetilde{p}p​ is constructed to emphasize exactly two cups while preserving the specified color and scene.

Before a revised prompt is used for training, UniEvo-VL applies a post-revision verification stage. The student model regenerates an image I′I'I′ from the revised prompt using the same noise:

I′=Mθgen(p~,ϵ).I' = \mathcal{M}_\theta^{\mathrm{gen}}(\widetilde{p}, \epsilon).I′=Mθgen​(p​,ϵ).

The regenerated image is then assessed by the understanding mode against the original prompt ppp. A revised prompt is accepted only if this second-round assessment is valid and accepts I′I'I′, indicating that the revised prompt leads to a better aligned image. The regenerated image I′I'I′ is used only for verifying the prompt and is not treated as a training objective later.

After collecting valid revised prompts, the method uses on-policy self-distillation. For a training triple (p,p~,ϵ)(p, \widetilde{p}, \epsilon)(p,p​,ϵ) from the current accepted set Ak\mathcal{A}_kAk​, the student policy performs a TTT-step denoising trajectory conditioned only on ppp:

τ={s0,…,sT},\tau = \{s_0, \dots, s_T\},τ={s0​,…,sT​},

where each sjs_jsj​ is an intermediate latent state. The student observes only the original prompt, matching the inference-time condition. The teacher observes the privileged prompt p~\widetilde{p}p​, which includes the corrective suggestion, and therefore produces richer denoising targets.

The training objective aligns the student one-step transition with the teacher transition at each sampled state. Let Mθgen(sj,p)\mathcal{M}_\theta^{\mathrm{gen}}(s_j, p)Mθgen​(sj​,p) denote the student transition and Mθˉgen(sj,p~)\mathcal{M}_{\bar\theta}^{\mathrm{gen}}(s_j, \widetilde{p})Mθˉgen​(sj​,p​) denote the teacher transition. The distillation loss is:

L(θ)=E(p,p~,ϵ)∼Ak, τ∼Mθgen(p,ϵ), sj∼τD(Mθgen(sg[sj],p), sg[Mθˉgen(sg[sj],p~)]),\mathcal{L}(\theta) = \mathbb{E}_{(p, \widetilde{p}, \epsilon) \sim \mathcal{A}_k,\, \tau \sim \mathcal{M}_\theta^{\mathrm{gen}}(p, \epsilon),\, s_j \sim \tau} D\left( \mathcal{M}_\theta^{\mathrm{gen}}(\mathrm{sg}[s_j], p), \, \mathrm{sg}\left[\mathcal{M}_{\bar\theta}^{\mathrm{gen}}(\mathrm{sg}[s_j], \widetilde{p})\right] \right),L(θ)=E(p,p​,ϵ)∼Ak​,τ∼Mθgen​(p,ϵ),sj​∼τ​D(Mθgen​(sg[sj​],p),sg[Mθˉgen​(sg[sj​],p​)]),

where D(⋅,⋅)D(\cdot, \cdot)D(⋅,⋅) is a divergence such as KL divergence or Jensen-Shannon divergence, and sg[⋅]\mathrm{sg}[\cdot]sg[⋅] denotes stop-gradient. Generation, assessment, and prompt synthesis are performed without gradient flow, so gradients update only the student parameters θ\thetaθ. In the flow-based implementation, the divergence is instantiated as squared error between deterministic latent transitions.

Experiment

The experiments evaluate a Qwen-Image generation and Qwen-VL feedback setting in which only generator LoRA parameters are updated, testing GenEval, GenEval2, and OCR text rendering with and without post-revision verification. Verified feedback broadly improves direct generation, with gains concentrated on prompts the base model initially struggles on and only small regressions on easier prompts. Training transfers reflection benefits while additional reflection remains useful, and stronger external critique yields better compositional results but its benefits depend on category-specific skills.

The table compares direct generation performance of Base, unverified UniEvo-VL, UniEvo-VL with GPT-5.6-Luna feedback, and verified UniEvo-VL across GenEval, GenEval2, and OCR. Verified UniEvo-VL exceeds Base on every reported metric across all three tasks, while the external critic configuration leads several GenEval and GenEval2 metrics. Task-specific differences remain, with verified feedback strongest on OCR and external feedback strongest on GenEval native and atomic metrics. Verified UniEvo-VL exceeds Base on every reported metric across GenEval, GenEval2, and OCR, though different configurations lead individual metrics, and the GPT-5.6-Luna external critic achieves the highest GenEval and GenEval2 native and atomic scores. On OCR, verified UniEvo-VL leads in native score and HumanPref, while unverified UniEvo-VL falls below Base on both measures; on GenEval2, unverified UniEvo-VL slightly trails Base on native Soft-TIFA despite improving atomic accuracy and HumanPref.

Training improves direct generation across all reported metrics on GenEval, GenEval2, and OCR. The evolved model's direct native scores approach or exceed the reflected base model on GenEval and OCR but remain below on GenEval2. Reflection further improves the evolved model on every metric, and the combined setting achieves the highest native scores across benchmarks. Direct generation improves after training on every reported metric across the three benchmarks. With direct generation, the evolved model approaches the reflected base model on GenEval and exceeds it on OCR, while remaining below it on GenEval2. Reflection gains after training shrink on GenEval but stay substantial on GenEval2 and change little on OCR. Combining training with reflection yields the highest native scores on all three benchmarks, though the reflected base model retains a slightly higher GenEval HumanPref score.

Stronger critic feedback improves compositional generation, but the gains are skill-dependent and uneven across benchmarks. The stronger critic reaches the highest overall GenEval score, with the largest improvements in spatial composition categories such as position and counting. On GenEval2, the stronger critic yields larger gains in two-object composition and color, while position and attribute scores improve only modestly. Stronger critic feedback produces the highest overall GenEval score and larger spatial composition gains, particularly in position and counting. On GenEval2, stronger critic feedback improves two-object composition and color more than Qwen feedback, while position and attribute gains remain limited. Native Soft-TIFA GM corroborates the trend: the stronger critic improves over baseline whereas Qwen feedback slightly declines.

These experiments evaluate text-to-image generation across GenEval, GenEval2, and OCR, comparing base models, trained UniEvo-VL variants, and critic-guided feedback. Verified feedback consistently improves over the base model on all three benchmarks, with external critic feedback leading on several GenEval metrics and verified feedback strongest on OCR. Training improves direct generation across benchmarks, and combining training with reflection achieves the highest native scores, although reflection gains shrink on GenEval and remain larger on GenEval2. Stronger critic feedback improves compositional generation, particularly spatial and two-object composition, but the gains are uneven and skill-dependent.


Créer de l'IA avec l'IA

De l'idée au lancement — accélérez votre développement IA avec le co-codage IA gratuit, un environnement prêt à l'emploi et le meilleur prix pour les GPU.

Codage assisté par IA
GPU prêts à l’emploi
Tarifs les plus avantageux

HyperAI Newsletters

Abonnez-vous à nos dernières mises à jour
Nous vous enverrons les dernières mises à jour de la semaine dans votre boîte de réception à neuf heures chaque lundi matin
Propulsé par MailChimp