HyperAIHyperAI

Command Palette

Search for a command to run...

UNIEVO-VL: マルチモーダルモデル自己改善のためのオンポリシー自己蒸留訓練手法

概要

現代のマルチモーダルモデルは、生成と理解を単一の統合システムへとまとめ、それによって自らのフィードバックを提供し、そこから学習することを可能にしている。この統合能力に着想を得て、我々は、テスト時計算の過程で建設的な自己修正フィードバックから学習する自己進化フレームワークUniEvo-VLを導入する。別個の、しばしばより大規模な教師に依存する代わりに、自己批評を特権情報として活用し、単一のマルチモーダルモデルに異なる文脈のもとで教師と生徒の両方の役割を担わせる。生徒は通常の質問のみを観測し、教師は特権的な批評に条件付けられる。学習では、生徒自身のサンプリング軌道上で、両者のノイズ除去拡散分布間の状態ごとのダイバージェンスを最小化する。実験により、UniEvo-VLはマルチモーダルモデルの画像生成能力を向上させると同時に、追加の内省情報に対する感度を維持することが示された。具体的には、オープンソースのQwen-image-2512を基盤として、GenEvalでは0.747から0.808へ、GenEval2 Soft-TIFAでは32.97から35.53へという顕著な性能向上が観測された。さらに、より強力な外部批評家(例:GPT5.6-Luna)を用いた試みは、優れた判定能力を持つマルチモーダルモデルがより高い自己進化の上限を見込めることを示している。最後に、テキストレンダリングに関する結果が混在していることから、自己改善は異なるタスク間で一様ではない可能性があることが示された。本研究は、外部の監督や誘導なしにマルチモーダルモデルを利用する際のユーザー体験を向上させるため、現在注目を集める再帰的自己改善の研究路線に光を当てることを目的とする。

One-sentence Summary

Stanford University, Johns Hopkins University, and colleagues introduce UniEvo-VL, a self-evolving framework that uses a single multimodal model as both teacher and student by conditioning the teacher on self-critiques and minimizing per-state divergence between denoising diffusion distributions, improving image generation in Qwen-image-2512 from 0.7470.7470.747 to 0.8080.8080.808 on GenEval without external supervision.

Key Contributions

  • UniEvo-VL introduces a critique-conditioned on-policy self-distillation framework in which one multimodal model acts as both teacher and student. The teacher conditions on self-generated corrective critiques while the student sees only the original prompt, and training matches denoising distributions along the student’s own sampling trajectories without corrected-image targets or reward-based policy optimization.
  • The work separates learned improvements from inference-time correction by comparing direct and reflection-assisted generation before and after training. This paired evaluation distinguishes gains retained in the initial model from the additional benefit of reflection at test time.
  • Experiments on Qwen-image-2512 improve GenEval from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53. Further analysis of critic choice and post-revision verification shows that attempts with more powerful external critics such as GPT5.6-Luna suggest a higher self-evolving ceiling, while gains are uneven across text-rendering tasks.

Introduction

Unified multimodal models can both generate and understand images, creating an opportunity for self-evolution through self-critique. In compositional image generation and visual text rendering, prior self-improvement methods rely on selecting self-generated outputs for fine-tuning, training with reflection triplets, or converting multimodal assessments into rewards. However, these methods do not directly transfer corrective feedback into the policy that generates from the original prompt along its own sampling trajectories; a critique states what should be fixed but not how intermediate denoising predictions should change. The authors address this gap with UniEvo-VL, a self-distillation framework in which a teacher model conditions on a critique-revised prompt while the student model observes only the original request, and their predictions are matched on the student’s own trajectories. This converts visual critique into dense state-wise supervision without corrected-image targets or reward-based policy optimization.

Method

UniEvo-VL is built around a single multimodal model Mθ\mathcal{M}_\thetaMθ​ that supports two operating modes: generation and understanding. In generation mode, the model maps a prompt ppp and sampled noise ϵ\epsilonϵ to an image:

I=Mθgen(p,ϵ).I = \mathcal{M}_\theta^{\mathrm{gen}}(p, \epsilon).I=Mθgen​(p,ϵ).

In understanding mode, the model evaluates an image III against a conditioning prompt ppp and returns a binary acceptance decision aaa together with a discrepancy description ccc:

(c,a)=Mθund(p,I).(c, a) = \mathcal{M}_\theta^{\mathrm{und}}(p, I).(c,a)=Mθund​(p,I).

When a=1a=1a=1, the image is considered aligned with the prompt and does not require revision. When a=0a=0a=0, the image is considered flawed, and ccc provides corrective feedback. Empty or unusable critique responses are treated as invalid and removed from the loop.

The authors connect generation and understanding through corrective conditioning. They define two policies. The student policy Mθ\mathcal{M}_\thetaMθ​ observes only the vanilla prompt ppp. The teacher policy is an exponential moving average version Mθˉ\mathcal{M}_{\bar\theta}Mθˉ​ of the same model, and it is allowed to observe privileged correction information. This setup enables the teacher to generate distillation targets that reflect the corrections while the student must learn to achieve the same behavior without access to the privileged context.

For an image that receives a valid rejection with a=0a=0a=0, the understanding mode synthesizes a privileged prompt p~\widetilde{p}p​ from the original prompt ppp and the corrective feedback ccc:

p~=Mθund(p,c).\widetilde{p} = \mathcal{M}_\theta^{\mathrm{und}}(p, c).p​=Mθund​(p,c).

The privileged prompt restates the requested scene while explicitly resolving the discrepancies identified in ccc. For example, if the original prompt requests two red cups but the model generates three cups, p~\widetilde{p}p​ is constructed to emphasize exactly two cups while preserving the specified color and scene.

Before a revised prompt is used for training, UniEvo-VL applies a post-revision verification stage. The student model regenerates an image I′I'I′ from the revised prompt using the same noise:

I′=Mθgen(p~,ϵ).I' = \mathcal{M}_\theta^{\mathrm{gen}}(\widetilde{p}, \epsilon).I′=Mθgen​(p​,ϵ).

The regenerated image is then assessed by the understanding mode against the original prompt ppp. A revised prompt is accepted only if this second-round assessment is valid and accepts I′I'I′, indicating that the revised prompt leads to a better aligned image. The regenerated image I′I'I′ is used only for verifying the prompt and is not treated as a training objective later.

After collecting valid revised prompts, the method uses on-policy self-distillation. For a training triple (p,p~,ϵ)(p, \widetilde{p}, \epsilon)(p,p​,ϵ) from the current accepted set Ak\mathcal{A}_kAk​, the student policy performs a TTT-step denoising trajectory conditioned only on ppp:

τ={s0,…,sT},\tau = \{s_0, \dots, s_T\},τ={s0​,…,sT​},

where each sjs_jsj​ is an intermediate latent state. The student observes only the original prompt, matching the inference-time condition. The teacher observes the privileged prompt p~\widetilde{p}p​, which includes the corrective suggestion, and therefore produces richer denoising targets.

The training objective aligns the student one-step transition with the teacher transition at each sampled state. Let Mθgen(sj,p)\mathcal{M}_\theta^{\mathrm{gen}}(s_j, p)Mθgen​(sj​,p) denote the student transition and Mθˉgen(sj,p~)\mathcal{M}_{\bar\theta}^{\mathrm{gen}}(s_j, \widetilde{p})Mθˉgen​(sj​,p​) denote the teacher transition. The distillation loss is:

L(θ)=E(p,p~,ϵ)∼Ak, τ∼Mθgen(p,ϵ), sj∼τD(Mθgen(sg[sj],p), sg[Mθˉgen(sg[sj],p~)]),\mathcal{L}(\theta) = \mathbb{E}_{(p, \widetilde{p}, \epsilon) \sim \mathcal{A}_k,\, \tau \sim \mathcal{M}_\theta^{\mathrm{gen}}(p, \epsilon),\, s_j \sim \tau} D\left( \mathcal{M}_\theta^{\mathrm{gen}}(\mathrm{sg}[s_j], p), \, \mathrm{sg}\left[\mathcal{M}_{\bar\theta}^{\mathrm{gen}}(\mathrm{sg}[s_j], \widetilde{p})\right] \right),L(θ)=E(p,p​,ϵ)∼Ak​,τ∼Mθgen​(p,ϵ),sj​∼τ​D(Mθgen​(sg[sj​],p),sg[Mθˉgen​(sg[sj​],p​)]),

where D(⋅,⋅)D(\cdot, \cdot)D(⋅,⋅) is a divergence such as KL divergence or Jensen-Shannon divergence, and sg[⋅]\mathrm{sg}[\cdot]sg[⋅] denotes stop-gradient. Generation, assessment, and prompt synthesis are performed without gradient flow, so gradients update only the student parameters θ\thetaθ. In the flow-based implementation, the divergence is instantiated as squared error between deterministic latent transitions.

Experiment

The experiments evaluate a Qwen-Image generation and Qwen-VL feedback setting in which only generator LoRA parameters are updated, testing GenEval, GenEval2, and OCR text rendering with and without post-revision verification. Verified feedback broadly improves direct generation, with gains concentrated on prompts the base model initially struggles on and only small regressions on easier prompts. Training transfers reflection benefits while additional reflection remains useful, and stronger external critique yields better compositional results but its benefits depend on category-specific skills.

The table compares direct generation performance of Base, unverified UniEvo-VL, UniEvo-VL with GPT-5.6-Luna feedback, and verified UniEvo-VL across GenEval, GenEval2, and OCR. Verified UniEvo-VL exceeds Base on every reported metric across all three tasks, while the external critic configuration leads several GenEval and GenEval2 metrics. Task-specific differences remain, with verified feedback strongest on OCR and external feedback strongest on GenEval native and atomic metrics. Verified UniEvo-VL exceeds Base on every reported metric across GenEval, GenEval2, and OCR, though different configurations lead individual metrics, and the GPT-5.6-Luna external critic achieves the highest GenEval and GenEval2 native and atomic scores. On OCR, verified UniEvo-VL leads in native score and HumanPref, while unverified UniEvo-VL falls below Base on both measures; on GenEval2, unverified UniEvo-VL slightly trails Base on native Soft-TIFA despite improving atomic accuracy and HumanPref.

Training improves direct generation across all reported metrics on GenEval, GenEval2, and OCR. The evolved model's direct native scores approach or exceed the reflected base model on GenEval and OCR but remain below on GenEval2. Reflection further improves the evolved model on every metric, and the combined setting achieves the highest native scores across benchmarks. Direct generation improves after training on every reported metric across the three benchmarks. With direct generation, the evolved model approaches the reflected base model on GenEval and exceeds it on OCR, while remaining below it on GenEval2. Reflection gains after training shrink on GenEval but stay substantial on GenEval2 and change little on OCR. Combining training with reflection yields the highest native scores on all three benchmarks, though the reflected base model retains a slightly higher GenEval HumanPref score.

Stronger critic feedback improves compositional generation, but the gains are skill-dependent and uneven across benchmarks. The stronger critic reaches the highest overall GenEval score, with the largest improvements in spatial composition categories such as position and counting. On GenEval2, the stronger critic yields larger gains in two-object composition and color, while position and attribute scores improve only modestly. Stronger critic feedback produces the highest overall GenEval score and larger spatial composition gains, particularly in position and counting. On GenEval2, stronger critic feedback improves two-object composition and color more than Qwen feedback, while position and attribute gains remain limited. Native Soft-TIFA GM corroborates the trend: the stronger critic improves over baseline whereas Qwen feedback slightly declines.

These experiments evaluate text-to-image generation across GenEval, GenEval2, and OCR, comparing base models, trained UniEvo-VL variants, and critic-guided feedback. Verified feedback consistently improves over the base model on all three benchmarks, with external critic feedback leading on several GenEval metrics and verified feedback strongest on OCR. Training improves direct generation across benchmarks, and combining training with reflection achieves the highest native scores, although reflection gains shrink on GenEval and remain larger on GenEval2. Stronger critic feedback improves compositional generation, particularly spatial and two-object composition, but the gains are uneven and skill-dependent.


AIでAIを構築

アイデアからローンチまで — 無料のAIコーディング支援、すぐに使える環境、最高のGPU価格でAI開発を加速。

AI コーディング補助
すぐに使える GPU
最適な料金体系

HyperAI Newsletters

最新情報を購読する
北京時間 毎週月曜日の午前9時 に、その週の最新情報をメールでお届けします
メール配信サービスは MailChimp によって提供されています