HyperAIHyperAI

Command Palette

Search for a command to run...

الفرض داخل التوزيع لتوليد مقاطع فيديو طويلة في وقت الاختبار

Jeongwoo Shin Youngyoon Choi Sangwoo Jo Hyunmog Kim Sungjoon Choi Joonseok Lee Jaewoong Choi Jaemoo Choi

الملخص

تتفوق نماذج الانتشار الفيديوي الانحدارية الذاتية (AR) الحديثة في توليد مقاطع فيديو قصيرة الأفق، غير أن توليد مقاطع فيديو طويلة ما يزال صعبًا بسبب الانجراف، حيث تتبدل الألوان والأنسجة وتضمحل ديناميكيات الحركة. تعتمد الأعمال الحالية أساسًا على تكييف KV، الذي ينتقي أو يعدّل مدخلات المفتاح والقيمة (KV) المخزنة مؤقتًا للحد من الانجراف. غير أننا نلاحظ أن تكييف KV وحده غير كافٍ، لأنه يفترض بقاء مدخلات KV المخزنة مؤقتًا داخل التوزيع. ويفشل هذا الافتراض خارج أفق التدريب: إذ لا يوجد ما يقيّد بناء مدخلات KV أثناء التوليد المتتابع، مما ينشئ مشكلة منشأ مدخلات KV التي تصبح فيها المدخلات المخزنة نفسها خارج التوزيع (OOD). ولمعالجة ذلك، نقترح الفرض داخل التوزيع (ID-Forcing)، وهو إطار عمل وقت الاختبار يعمل على مواءمة كلٍّ من تخزين KV المؤقت وتكييف KV مع إعدادات التدريب. وتمنع آليته الأساسية، selfcaching، نشوء مدخلات KV خارج التوزيع من مصدرها؛ إذ تُخزَّن كل قطعة مؤقتًا دون الانتباه إلى مدخلات KV السابقة، مما يُبقي النافذة المنزلقة داخل التوزيع تمامًا. وبناءً عليه، يوسّع ID-Forcing بسلاسة النماذج قصيرة الأفق لتوليد فيديو على مدى دقائق. وتُظهر تقييمات موسّعة أن طريقتنا تظل قادرة على المنافسة في معيار قياسي لتوليد الفيديو، بينما تتفوق بوضوح على الأعمال السابقة في الحد من الانجراف، وفق ما تؤكده كلٌّ من مقاييس الانجراف لدينا ودراسة المستخدمين.

One-sentence Summary

Researchers from Seoul National University, Korea University, Sungkyunkwan University, and Georgia Institute of Technology propose In-Distribution Forcing (ID-Forcing), a test-time framework that aligns both KV caching and KV conditioning with training configurations and uses a selfcaching mechanism to prevent out-of-distribution key-value entries at their source, outperforming prior KV conditioning methods in mitigating drifting and extending short-horizon video diffusion models to minute-scale generation.

Key Contributions

  • The paper identifies the KV-provenance problem in autoregressive video diffusion: beyond the training horizon, cached key-value entries from rolling-window generation become out-of-distribution, so KV conditioning alone cannot prevent drifting.
  • It introduces a training-free test-time framework, In-Distribution Forcing (ID-Forcing), whose self-caching mechanism caches each chunk without attending to prior KV entries; at the conditioning level it bounds the window length to L ≤ N−1 and permanently fixes the first chunk κ0.
  • Evaluations across durations, backbones, and metrics show that ID-Forcing remains competitive on a standard video generation benchmark and substantially outperforms prior test-time approaches in mitigating drifting, as measured by drift metrics and a user study, including minute-scale generation.

Introduction

Autoregressive video diffusion enables long-horizon synthesis by generating video chunks sequentially and conditioning each chunk on key-value cache entries from earlier chunks, which matters for streaming, world models, and long-form storytelling. A key limitation is drifting: beyond the finite training horizon, colors and textures shift and motion decays because test-time rollouts enter conditioning configurations not seen during training. Prior methods mostly adjust KV conditioning, such as attention sinks or temporal position embeddings, but they assume the cached KV entries themselves remain in-distribution. The authors show this assumption can fail because each KV entry is built from its own attended context, called its provenance; beyond the training horizon, caching itself uses unseen provenance. Their main contribution is In-Distribution Forcing, which uses self-caching and a bounded rolling window so that both KV caching and KV conditioning stay within training-time configurations, allowing a short video model to generate minute-scale video with reduced drift.

Method

The authors propose In-Distribution Forcing (ID-Forcing), a training-free framework designed to keep test-time extrapolation strictly within the training distribution. Unlike prior works that treat drifting as a consequence of discarding informative past context and only modify KV conditioning, ID-Forcing views drifting as a deviation from the training distribution that occurs in both KV caching and KV conditioning. Therefore, the framework introduces two-level control over KV operations to prevent out-of-distribution states from accumulating.

Level 1: KV Caching and the KV-Provenance Problem

The authors identify that the same clean chunk maps to different points in KV space depending on its provenance, which refers to the existing KV entries within the caching window. When this provenance involves unseen context beyond the training horizon, the resulting KV entry falls out-of-distribution, a phenomenon termed the KV-provenance problem.

As illustrated in the framework diagram, generating a chunk from an out-of-distribution provenance causes subsequent generations to inherit this state, leading to progressive drifting. To resolve this, the authors introduce Self-Caching, a rule that restricts every provenance to configurations executed during training. For a conditioning window of length LLL, the first ℓ\ellℓ entries are cached by attending only to themselves, denoted as E(xj;∅)E(x_j; \varnothing)E(xj​;∅). The remaining L−ℓL - \ellL−ℓ entries are cached autoregressively on top of the last self-cached entry. This ensures that no KV entry is ever written under an unseen context, keeping the provenance strictly in-distribution throughout extrapolation. The caching rule is formalized as:

κj={E(xj;∅),i−L≤j<i−L+ℓ,E(xj;κi−L+ℓ−1:j−1),i−L+ℓ≤j≤i−1.\kappa_{j} = \left\{ \begin{array}{l l} E (x_{j}; \varnothing), & i - L \leq j < i - L + \ell, \\ E (x_{j}; \kappa_{i - L + \ell - 1: j - 1}), & i - L + \ell \leq j \leq i - 1. \end{array} \right.κj​={E(xj​;∅),E(xj​;κi−L+ℓ−1:j−1​),​i−L≤j<i−L+ℓ,i−L+ℓ≤j≤i−1.​

Level 2: KV Conditioning and Exact Rolling Window

While Self-Caching ensures stored KV entries are in-distribution, the authors note that how these entries are read must also match training. They address KV conditioning by treating the first chunk as a persistent condition and enabling an exact rolling window.

During training, every generation is conditioned on the first chunk KV entry, κ0\kappa_0κ0​. Because the causal 3D-VAE encodes the initial latent frame from a single pixel frame, κ0\kappa_0κ0​ is distributionally unique and irreplaceable. Evicting it during extrapolation would itself be out-of-distribution. Therefore, ID-Forcing retains κ0\kappa_0κ0​ at slot 0 as a persistent sink, differing from standard attention sinks that retain multiple early entries for empirical stability.

As shown in the figure below, the conditioning window is constrained to a length L<NL < NL<N, holding κ0\kappa_0κ0​ at slot 0 and the L−1L - 1L−1 most recent KV entries in the remaining slots. This strictly bounds the total window size within the training budget. After generating each chunk, the oldest non-sink KV entry is evicted, and the remaining entries are re-rotated by R−fR_{-f}R−f​, while κ0\kappa_0κ0​ remains fixed. The conditioning rule is expressed as:

{κ0}∪{κi−L+1,…,κi−1}←{κ0}∪R−f[{κi−L+1,…,κi−1}].\{\kappa_{0}\} \cup \{\kappa_{i - L + 1}, \dots, \kappa_{i - 1}\} \leftarrow \{\kappa_{0}\} \cup R_{-f} \big[ \{\kappa_{i - L + 1}, \dots, \kappa_{i - 1}\} \big].{κ0​}∪{κi−L+1​,…,κi−1​}←{κ0​}∪R−f​[{κi−L+1​,…,κi−1​}].

Because Self-Caching ensures that every cached KV entry possesses no external provenance to lose during eviction, this rolling window remains exactly in-distribution, matching the temporal distances observed during training.

Experiment

ID-Forcing is evaluated for minute-scale video generation beyond the training horizon on Self-Forcing and LongLive base models, using 120- and 240-second videos and comparing against Deep Forcing, ∞-RoPE, and MemRoPE with automated drift metrics and a user study. The main results show that ID-Forcing reduces color and motion drift while preserving subject consistency, visual quality, and dynamic motion, and qualitative comparisons confirm it avoids the collapse and artifacts seen in baselines. Ablations validate the contributions of self-caching, retaining the first chunk as a sink, and re-rotating the conditioning window, showing that these components together suppress drifting and that self-caching is especially important for stable long-horizon generation.

ID-Forcing shows the strongest long-duration stability among the compared methods, with the highest Dynamic Degree and clear gains on color and motion drift metrics in the reported settings. It also remains competitive on Aesthetic Quality, Imaging Quality, and Subject Consistency, while Motion Smoothness is lower primarily because that metric favors less dynamic outputs. Qualitative comparisons further indicate that ID-Forcing preserves subject identity, color fidelity, and motion over long videos. ID-Forcing achieves the highest Dynamic Degree in the reported settings, while Self-Forcing and LongLive baselines tend toward frame freezing at longer durations. It also shows substantial gains on color and motion drift metrics and remains competitive on quality and subject consistency, with Motion Smoothness as the only metric favoring less dynamic outputs.

In a pairwise user study, participants preferred ID-Forcing over all four Self-Forcing-based baselines across every evaluated dimension. The advantage was strongest against Self-Forcing, where preferences were near unanimous, including overall preference. Against MemRoPE, the closest dimension was dynamic degree, but ID-Forcing still received majority preference. ID-Forcing was preferred over every baseline on color consistency, background consistency, subject consistency, dynamic degree, temporal flicker, and overall preference. The strongest preference gap appeared against Self-Forcing, with near-unanimous ratings across most dimensions, while the smallest lead was on dynamic degree against MemRoPE.

Experiments evaluate ID-Forcing against Self-Forcing and LongLive baselines on long-duration video generation, using automated metrics, qualitative comparisons, and a pairwise user study. ID-Forcing achieves the strongest long-duration stability and highest Dynamic Degree, with clear gains on color and motion drift while remaining competitive on quality and subject consistency; Motion Smoothness is lower mainly because that metric favors less dynamic outputs. Qualitative results show preserved subject identity, color fidelity, and motion over long videos. In the user study, participants preferred ID-Forcing over all four Self-Forcing-based baselines across every evaluated dimension, including overall preference, with near-unanimous preference against Self-Forcing and the smallest lead on dynamic degree against MemRoPE.


بناء الذكاء الاصطناعي بالذكاء الاصطناعي

من الفكرة إلى الإطلاق — سرّع تطوير الذكاء الاصطناعي الخاص بك مع المساعدة البرمجية المجانية بالذكاء الاصطناعي، وبيئة جاهزة للاستخدام، وأفضل أسعار لوحدات معالجة الرسومات.

البرمجة التعاونية باستخدام الذكاء الاصطناعي
وحدات GPU جاهزة للعمل
أفضل الأسعار

HyperAI Newsletters

اشترك في آخر تحديثاتنا
سنرسل لك أحدث التحديثات الأسبوعية إلى بريدك الإلكتروني في الساعة التاسعة من صباح كل يوم اثنين
مدعوم بواسطة MailChimp