Command Palette
Search for a command to run...
UNIEVO-VL: وصفة تدريب للتقطير الذاتي وفق السياسة من أجل التحسين الذاتي للنماذج متعددة الوسائط
UNIEVO-VL: وصفة تدريب للتقطير الذاتي وفق السياسة من أجل التحسين الذاتي للنماذج متعددة الوسائط
الملخص
تجمع النماذج متعددة الوسائط الحديثة بين التوليد والفهم في نظام موحد واحد، مما يمكّنها من تقديم تغذية راجعة ذاتية والتعلم منها. وانطلاقًا من هذه القدرة الموحدة، نقدم UniEvo-VL، وهو إطار ذاتي التطور يمكّن النماذج متعددة الوسائط من التعلم من هذه التغذية الراجعة البناءة الناتجة عن التصحيح الذاتي أثناء الحوسبة في زمن الاختبار. فبدلاً من الاعتماد على معلم منفصل وغالبًا أكبر حجمًا، نستغل النقد الذاتي للنموذج بوصفه معلومات مميزة، ونطلب من نموذج واحد متعدد الوسائط أن يؤدي دور المعلم والطالب في آن واحد بسياقات مختلفة؛ إذ لا يرى الطالب سوى السؤال المجرد، بينما يحصل المعلم على النقد المميز بوصفه سياقًا. ثم يقلل التدريب التباعد بين توزيعات الانتشار المزيلة للضوضاء لكل حالة عبر مسارات أخذ العينات الخاصة بالطالب نفسه. تُظهر التجارب أن UniEvo-VL يحسن قدرات توليد الصور لدى النماذج متعددة الوسائط، مع الحفاظ على حساسيتها لمعلومات التأمل الإضافية. وبالتحديد، نبني عملنا على النموذج مفتوح المصدر Qwen-image-2512 ونلاحظ تحسنًا كبيرًا في الأداء من 0.747 إلى 0.808 على GenEval، ومن 32.97 إلى 35.53 على GenEval2 Soft-TIFA. علاوة على ذلك، تُظهر المحاولات مع نقاد خارجيين أقوى (مثل GPT5.6-Luna) أن النماذج متعددة الوسائط ذات القدرات القوية على التحكيم يمكنها توقع سقف أعلى للتطور الذاتي. وأخيرًا وليس آخرًا، تُظهر نتائج عرض النصوص المختلطة أن تحسيناتنا الذاتية قد لا تكون موحدة عبر المهام المختلفة. تهدف دراستنا إلى تسليط الضوء على خط البحث الحيوي الراهن في التحسين الذاتي التكراري لتعزيز تجربة المستخدم عند استخدام النماذج متعددة الوسائط دون إشراف خارجي أو توجيه.
One-sentence Summary
Stanford University, Johns Hopkins University, and colleagues introduce UniEvo-VL, a self-evolving framework that uses a single multimodal model as both teacher and student by conditioning the teacher on self-critiques and minimizing per-state divergence between denoising diffusion distributions, improving image generation in Qwen-image-2512 from 0.747 to 0.808 on GenEval without external supervision.
Key Contributions
- UniEvo-VL introduces a critique-conditioned on-policy self-distillation framework in which one multimodal model acts as both teacher and student. The teacher conditions on self-generated corrective critiques while the student sees only the original prompt, and training matches denoising distributions along the student’s own sampling trajectories without corrected-image targets or reward-based policy optimization.
- The work separates learned improvements from inference-time correction by comparing direct and reflection-assisted generation before and after training. This paired evaluation distinguishes gains retained in the initial model from the additional benefit of reflection at test time.
- Experiments on Qwen-image-2512 improve GenEval from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53. Further analysis of critic choice and post-revision verification shows that attempts with more powerful external critics such as GPT5.6-Luna suggest a higher self-evolving ceiling, while gains are uneven across text-rendering tasks.
Introduction
Unified multimodal models can both generate and understand images, creating an opportunity for self-evolution through self-critique. In compositional image generation and visual text rendering, prior self-improvement methods rely on selecting self-generated outputs for fine-tuning, training with reflection triplets, or converting multimodal assessments into rewards. However, these methods do not directly transfer corrective feedback into the policy that generates from the original prompt along its own sampling trajectories; a critique states what should be fixed but not how intermediate denoising predictions should change. The authors address this gap with UniEvo-VL, a self-distillation framework in which a teacher model conditions on a critique-revised prompt while the student model observes only the original request, and their predictions are matched on the student’s own trajectories. This converts visual critique into dense state-wise supervision without corrected-image targets or reward-based policy optimization.
Method
UniEvo-VL is built around a single multimodal model Mθ that supports two operating modes: generation and understanding. In generation mode, the model maps a prompt p and sampled noise ϵ to an image:
I=Mθgen(p,ϵ).In understanding mode, the model evaluates an image I against a conditioning prompt p and returns a binary acceptance decision a together with a discrepancy description c:
(c,a)=Mθund(p,I).When a=1, the image is considered aligned with the prompt and does not require revision. When a=0, the image is considered flawed, and c provides corrective feedback. Empty or unusable critique responses are treated as invalid and removed from the loop.
The authors connect generation and understanding through corrective conditioning. They define two policies. The student policy Mθ observes only the vanilla prompt p. The teacher policy is an exponential moving average version Mθˉ of the same model, and it is allowed to observe privileged correction information. This setup enables the teacher to generate distillation targets that reflect the corrections while the student must learn to achieve the same behavior without access to the privileged context.
For an image that receives a valid rejection with a=0, the understanding mode synthesizes a privileged prompt p from the original prompt p and the corrective feedback c:
p=Mθund(p,c).The privileged prompt restates the requested scene while explicitly resolving the discrepancies identified in c. For example, if the original prompt requests two red cups but the model generates three cups, p is constructed to emphasize exactly two cups while preserving the specified color and scene.
Before a revised prompt is used for training, UniEvo-VL applies a post-revision verification stage. The student model regenerates an image I′ from the revised prompt using the same noise:
I′=Mθgen(p,ϵ).The regenerated image is then assessed by the understanding mode against the original prompt p. A revised prompt is accepted only if this second-round assessment is valid and accepts I′, indicating that the revised prompt leads to a better aligned image. The regenerated image I′ is used only for verifying the prompt and is not treated as a training objective later.
After collecting valid revised prompts, the method uses on-policy self-distillation. For a training triple (p,p,ϵ) from the current accepted set Ak, the student policy performs a T-step denoising trajectory conditioned only on p:
τ={s0,…,sT},where each sj is an intermediate latent state. The student observes only the original prompt, matching the inference-time condition. The teacher observes the privileged prompt p, which includes the corrective suggestion, and therefore produces richer denoising targets.
The training objective aligns the student one-step transition with the teacher transition at each sampled state. Let Mθgen(sj,p) denote the student transition and Mθˉgen(sj,p) denote the teacher transition. The distillation loss is:
L(θ)=E(p,p,ϵ)∼Ak,τ∼Mθgen(p,ϵ),sj∼τD(Mθgen(sg[sj],p),sg[Mθˉgen(sg[sj],p)]),where D(⋅,⋅) is a divergence such as KL divergence or Jensen-Shannon divergence, and sg[⋅] denotes stop-gradient. Generation, assessment, and prompt synthesis are performed without gradient flow, so gradients update only the student parameters θ. In the flow-based implementation, the divergence is instantiated as squared error between deterministic latent transitions.
Experiment
The experiments evaluate a Qwen-Image generation and Qwen-VL feedback setting in which only generator LoRA parameters are updated, testing GenEval, GenEval2, and OCR text rendering with and without post-revision verification. Verified feedback broadly improves direct generation, with gains concentrated on prompts the base model initially struggles on and only small regressions on easier prompts. Training transfers reflection benefits while additional reflection remains useful, and stronger external critique yields better compositional results but its benefits depend on category-specific skills.
The table compares direct generation performance of Base, unverified UniEvo-VL, UniEvo-VL with GPT-5.6-Luna feedback, and verified UniEvo-VL across GenEval, GenEval2, and OCR. Verified UniEvo-VL exceeds Base on every reported metric across all three tasks, while the external critic configuration leads several GenEval and GenEval2 metrics. Task-specific differences remain, with verified feedback strongest on OCR and external feedback strongest on GenEval native and atomic metrics. Verified UniEvo-VL exceeds Base on every reported metric across GenEval, GenEval2, and OCR, though different configurations lead individual metrics, and the GPT-5.6-Luna external critic achieves the highest GenEval and GenEval2 native and atomic scores. On OCR, verified UniEvo-VL leads in native score and HumanPref, while unverified UniEvo-VL falls below Base on both measures; on GenEval2, unverified UniEvo-VL slightly trails Base on native Soft-TIFA despite improving atomic accuracy and HumanPref.
Training improves direct generation across all reported metrics on GenEval, GenEval2, and OCR. The evolved model's direct native scores approach or exceed the reflected base model on GenEval and OCR but remain below on GenEval2. Reflection further improves the evolved model on every metric, and the combined setting achieves the highest native scores across benchmarks. Direct generation improves after training on every reported metric across the three benchmarks. With direct generation, the evolved model approaches the reflected base model on GenEval and exceeds it on OCR, while remaining below it on GenEval2. Reflection gains after training shrink on GenEval but stay substantial on GenEval2 and change little on OCR. Combining training with reflection yields the highest native scores on all three benchmarks, though the reflected base model retains a slightly higher GenEval HumanPref score.
Stronger critic feedback improves compositional generation, but the gains are skill-dependent and uneven across benchmarks. The stronger critic reaches the highest overall GenEval score, with the largest improvements in spatial composition categories such as position and counting. On GenEval2, the stronger critic yields larger gains in two-object composition and color, while position and attribute scores improve only modestly. Stronger critic feedback produces the highest overall GenEval score and larger spatial composition gains, particularly in position and counting. On GenEval2, stronger critic feedback improves two-object composition and color more than Qwen feedback, while position and attribute gains remain limited. Native Soft-TIFA GM corroborates the trend: the stronger critic improves over baseline whereas Qwen feedback slightly declines.
These experiments evaluate text-to-image generation across GenEval, GenEval2, and OCR, comparing base models, trained UniEvo-VL variants, and critic-guided feedback. Verified feedback consistently improves over the base model on all three benchmarks, with external critic feedback leading on several GenEval metrics and verified feedback strongest on OCR. Training improves direct generation across benchmarks, and combining training with reflection achieves the highest native scores, although reflection gains shrink on GenEval and remain larger on GenEval2. Stronger critic feedback improves compositional generation, particularly spatial and two-object composition, but the gains are uneven and skill-dependent.