HyperAIHyperAI

Command Palette

Search for a command to run...

DuoMatching: 少数ステップ動画生成のための同時分布・周辺分布マッチング

Jiahao Zhan Yan Wang Yongrui Ma Qunliang Xing Ruchang Yao Runtao Liu Shijie Zhao Tianfan Xue

概要

ストリーミング動画生成は,動画フレームの同時分布を実動画分布に対する動画教師の近似に一致させる分布マッチング蒸留(DMD)の恩恵を受けてきた。この同時分布マッチングは自己回帰的ロールアウト中のドリフトを軽減するが,視覚品質と意味的整合性には依然として限界が残る。この限界に対処するため,我々は,同時分布と周辺分布を統合した定式化により実動画分布を近似する分布マッチングフレームワークDuoMatchingを提案する。既存の同時分布マッチングの定式化に加え,追加の周辺分布マッチング目的は,画像生成器から専用のフレーム単位教師信号を提供し,画像生成器に由来する相補的な視覚的・意味的事前知識を伝達する。このフレーム単位教師信号を動画生成へ適用するために,我々は動画の生徒モデルと画像の教師モデルの間の潜在表現の不一致を解消するLatentBridgeを導入する。Latent Variation Samplingは,このようなフレーム単位教師信号を異なる時間セグメントへさらに分散させ,冗長性を低減する。実験により,DuoMatchingは動きのダイナミクスを概ね保持しながら視覚品質・構図・意味的整合性を向上させることが示された。人間による評価では,評価したすべてのベースラインに対して80%を超える総合選好率を示した。

One-sentence Summary

Researchers from MMLab, CUHK, ByteDance Inc., HKUST, and CPII under InnoHK propose DuoMatching, a joint-marginal distribution matching framework for few-step video generation that adds frame-level image-teacher supervision via LatentBridge to resolve latent representation mismatches and via Latent Variation Sampling to reduce temporal redundancy, thereby improving visual quality, composition, and semantic alignment while preserving motion dynamics and achieving over 80%80\%80% human preference against evaluated baselines.

Key Contributions

  • DuoMatching is a unified joint-marginal distribution matching framework for few-step video generation that adds frame-level supervision from an image teacher to existing joint video distribution matching.
  • Within DuoMatching, LatentBridge resolves the latent representation mismatch between the video student and image teacher, while Latent Variation Sampling distributes marginal supervision across distinct temporal segments to reduce redundancy.
  • Experiments demonstrate improvements in fine-grained visual quality, visual composition, and semantic alignment in causal and bidirectional generation while largely preserving temporal quality, with human evaluations reporting overall preference rates above 80% against all evaluated baselines.

Introduction

Recent video generation models achieve strong visual quality and temporal coherence but require many sequential denoising steps, limiting real-time applications such as interactive world simulation and digital entertainment. Distribution matching distillation compresses multi-step diffusion or flow models into few-step generators, but existing joint DMD methods mainly match the full video distribution and do not explicitly supervise individual frame distributions. This leaves weaknesses in fine-grained appearance and semantic alignment. The authors propose DuoMatching, a unified joint-marginal distribution matching framework that combines joint DMD with a video teacher and marginal DMD with an image teacher to strengthen frame-level supervision. They further introduce LatentBridge to resolve latent space mismatches for temporally compressed video latents and Latent Variation Sampling to reduce redundant supervision. Experiments show improved visual quality, composition, and semantic alignment while largely preserving temporal dynamics.

Method

The authors formulate video generation as a distribution matching problem, aiming to align the generator's output distribution with the real data distribution. Since the score of the real distribution is unavailable, prior work approximates it using a pretrained video teacher, resulting in a joint distribution matching objective. In this setup, the score for each latent slice is conditioned on all other slices. Consequently, improvements to an individual frame might be suppressed if they introduce inconsistencies with surrounding frames. As shown in the figure below, relying solely on joint matching leads to limitations in visual quality and semantic alignment, such as missing fine-grained details in generated frames.

To overcome these limitations, the authors propose DuoMatching, a unified framework that reformulates distribution matching through both joint and marginal perspectives. The framework retains the video teacher to constrain cross-frame dependencies via joint matching, while introducing an image teacher to provide direct frame-level marginal supervision. The combined objective balances the joint distribution estimate from the video teacher and the complementary marginal estimate from the image teacher.

A significant challenge in leveraging image teachers is the structural mismatch between video and image latent spaces. Video VAEs typically employ temporal compression, mapping a single latent slice to multiple RGB frames, whereas image VAEs map one latent to a single frame. To resolve this, the authors introduce LatentBridge, a lightweight differentiable module. As illustrated in the framework diagram, LatentBridge maps a video latent slice, conditioned on the preceding slice and a specific local frame index, to a frame-specific representation within the image teacher's latent space. The module is pretrained using paired video and image latents with an L1 reconstruction loss, enabling gradients from the marginal matching objective to propagate efficiently to the student generator.

To further optimize the marginal matching process, the authors design Latent Variation Sampling. Supervising randomly sampled frames can lead to redundant computations in slowly changing regions, potentially suppressing motion dynamics. Instead, this sampling strategy computes the mean squared difference between adjacent clean video latent slices to measure temporal variation. The sequence is partitioned into contiguous segments based on the positions with the highest variation. By uniformly sampling one latent slice from each segment, the method ensures that the limited supervision budget is distributed across diverse temporal regions, thereby improving temporal coverage and reducing redundancy in frame-level supervision.

Experiment

The experiments evaluate causal and bidirectional video generation models trained with DMD, comparing DuoMatching against autoregressive and bidirectional baselines under matched inference budgets. Results on VBench and human preference tests show DuoMatching improves visual quality, semantic alignment, and composition while preserving motion quality, with no additional inference cost. Ablations confirm that explicit marginal matching, stronger image teacher priors such as Qwen-Image, and LatentBridge contribute to these gains, and that latent variation sampling allocates frame-level supervision more effectively than uniform or stratified sampling.

DuoMatching improves overall VBench performance over matched baselines in both bidirectional full-video generation and causal autoregressive generation. In bidirectional mode, gains appear in semantic, aesthetic, imaging, and dynamic scores while smoothness remains nearly unchanged. In causal mode, it improves visual quality and semantic alignment while largely preserving motion quality, without additional inference computation. In bidirectional full-video generation with matched NFE, DuoMatching outperforms CausVid in total score and shows gains in semantic, aesthetic, imaging, and dynamic metrics, with smoothness similar. In causal autoregressive generation, DuoMatching achieves higher video generation quality across one-step, two-step, and four-step settings, with notable gains in visual quality and semantic alignment while motion quality is largely preserved. All reported gains are obtained without increasing inference computation, preserving the real-time generation capability of the causal baseline.

DuoMatching was preferred over all baselines in overall quality, with overall preference rates above 80 percent in every pairing. Its advantages were concentrated in visual quality and semantic alignment, while temporal and motion preferences were near or above parity, indicating that motion quality was largely preserved. Overall preference rates favored DuoMatching against every baseline by a clear margin. Visual quality and semantic alignment preferences consistently favored DuoMatching, whereas temporal and motion preferences remained near or above parity.

Using Wan2.1-14B as an image teacher improves all reported video student metrics over the no-teacher baseline. Across SDXL, FLUX.2-4B, and Qwen-Image, higher teacher HPSv3 scores correspond to consistent gains in semantic, aesthetic, and imaging scores, while dynamic degree and motion smoothness remain relatively stable. Qwen-Image achieves the strongest semantic, aesthetic, imaging, and total results among the compared image teachers. Wan2.1-14B as an image teacher improves performance across all reported metrics compared with the no-teacher baseline. For SDXL, FLUX.2-4B, and Qwen-Image, higher teacher HPSv3 scores align with consistent improvements in semantic, aesthetic, and imaging scores, while dynamic degree and motion smoothness stay relatively stable.

Enabling marginal DMD with direct latent mapping improves semantic and visual quality but substantially lowers dynamic quality. LatentBridge restores dynamic quality while improving semantic and imaging scores further, with only a small memory increase over the direct baseline. Decode-Encode was infeasible under the training configuration due to out-of-memory. Applying marginal DMD directly to temporally compressed video latents improves semantic and imaging scores but reduces Dynamic score compared with the no-marginal-DMD baseline. LatentBridge yields the best semantic and imaging results among feasible configurations while largely retaining dynamic quality. Decode-Encode runs out of memory, whereas LatentBridge adds only a small peak memory overhead over the Direct baseline.

Increasing the number of sampled latent slices from 2 to 4 improves total scores for uniform random sampling and Latent Variation Sampling, but further increasing to 8 reduces both dynamic and total scores. Latent Variation Sampling consistently outperforms uniform random sampling across all tested slice counts and also exceeds temporally stratified sampling at four slices. The results indicate that a moderate slice count with Latent Variation Sampling better balances temporal quality and overall performance while limiting computational cost. Latent Variation Sampling achieves the highest dynamic and total scores at K=4 among all tested sampling strategies. Raising K from 2 to 4 improves total scores, whereas raising K to 8 lowers both dynamic and total scores. Uniform random sampling trails Latent Variation Sampling at every tested K. At K=4, temporally stratified sampling scores between uniform random sampling and Latent Variation Sampling on dynamic and total metrics.

The experiments evaluate DuoMatching in bidirectional full-video and causal autoregressive generation, showing consistent gains in visual quality and semantic alignment while preserving temporal smoothness and inference cost, with human evaluations confirming clear overall preference and motion parity. Additional ablations validate that using Wan2.1-14B as an image teacher improves student metrics across baselines, and that replacing direct marginal DMD with LatentBridge restores dynamic quality while improving semantic and imaging scores with only modest memory overhead. Finally, a moderate number of latent slices under Latent Variation Sampling, especially four slices, provides the best balance between temporal quality and overall performance compared with uniform or stratified sampling.


AIでAIを構築

アイデアからローンチまで — 無料のAIコーディング支援、すぐに使える環境、最高のGPU価格でAI開発を加速。

AI コーディング補助
すぐに使える GPU
最適な料金体系

HyperAI Newsletters

最新情報を購読する
北京時間 毎週月曜日の午前9時 に、その週の最新情報をメールでお届けします
メール配信サービスは MailChimp によって提供されています