Command Palette
Search for a command to run...
テスト時における長尺動画生成のためのIn-Distribution Forcing
テスト時における長尺動画生成のためのIn-Distribution Forcing
Jeongwoo Shin Youngyoon Choi Sangwoo Jo Hyunmog Kim Sungjoon Choi Joonseok Lee Jaewoong Choi Jaemoo Choi
概要
現代の自己回帰型(AR)動画拡散モデルは短区間の動画生成には優れるが、長尺動画の生成では、色やテクスチャが変化し運動のダイナミクスが失われるドリフトが原因で依然として困難である。既存研究は主にKV条件付けに依存し、キャッシュされたキー・バリュー(KV)エントリを選択または修正することでドリフトを軽減する。しかし我々は、KV条件付けだけでは不十分であることを観察した。これは、キャッシュされたKVエントリが分布内に留まるという仮定に依拠しているためである。この仮定は学習ホライズンを超えると成り立たない。ロールアウト中にKVエントリの構築を制約するものは何もなく、キャッシュエントリ自体が分布外(OOD)になるというKV来歴問題が生じる。これに対処するため、我々はIn-Distribution Forcing(ID-Forcing)を提案する。これはKVキャッシングとKV条件付けの両方を学習時の設定に整合させるテスト時フレームワークである。その中核機構であるselfcachingは、OODなKVエントリをその発生源で防ぐ。各チャンクは過去のKVエントリにアテンションせずにキャッシュされ、ローリングウィンドウを正確に分布内に保つ。結果として、ID-Forcingは短区間モデルを分単位の動画生成へシームレスに拡張する。広範な評価により、本手法は標準的な動画生成ベンチマークで競争力を維持しながら、ドリフトの軽減において先行研究を大幅に上回ることが示され、これは我々のドリフト指標とユーザスタディの両方によって検証された。
One-sentence Summary
Researchers from Seoul National University, Korea University, Sungkyunkwan University, and Georgia Institute of Technology propose In-Distribution Forcing (ID-Forcing), a test-time framework that aligns both KV caching and KV conditioning with training configurations and uses a selfcaching mechanism to prevent out-of-distribution key-value entries at their source, outperforming prior KV conditioning methods in mitigating drifting and extending short-horizon video diffusion models to minute-scale generation.
Key Contributions
- The paper identifies the KV-provenance problem in autoregressive video diffusion: beyond the training horizon, cached key-value entries from rolling-window generation become out-of-distribution, so KV conditioning alone cannot prevent drifting.
- It introduces a training-free test-time framework, In-Distribution Forcing (ID-Forcing), whose self-caching mechanism caches each chunk without attending to prior KV entries; at the conditioning level it bounds the window length to L ≤ N−1 and permanently fixes the first chunk κ0.
- Evaluations across durations, backbones, and metrics show that ID-Forcing remains competitive on a standard video generation benchmark and substantially outperforms prior test-time approaches in mitigating drifting, as measured by drift metrics and a user study, including minute-scale generation.
Introduction
Autoregressive video diffusion enables long-horizon synthesis by generating video chunks sequentially and conditioning each chunk on key-value cache entries from earlier chunks, which matters for streaming, world models, and long-form storytelling. A key limitation is drifting: beyond the finite training horizon, colors and textures shift and motion decays because test-time rollouts enter conditioning configurations not seen during training. Prior methods mostly adjust KV conditioning, such as attention sinks or temporal position embeddings, but they assume the cached KV entries themselves remain in-distribution. The authors show this assumption can fail because each KV entry is built from its own attended context, called its provenance; beyond the training horizon, caching itself uses unseen provenance. Their main contribution is In-Distribution Forcing, which uses self-caching and a bounded rolling window so that both KV caching and KV conditioning stay within training-time configurations, allowing a short video model to generate minute-scale video with reduced drift.
Method
The authors propose In-Distribution Forcing (ID-Forcing), a training-free framework designed to keep test-time extrapolation strictly within the training distribution. Unlike prior works that treat drifting as a consequence of discarding informative past context and only modify KV conditioning, ID-Forcing views drifting as a deviation from the training distribution that occurs in both KV caching and KV conditioning. Therefore, the framework introduces two-level control over KV operations to prevent out-of-distribution states from accumulating.
Level 1: KV Caching and the KV-Provenance Problem
The authors identify that the same clean chunk maps to different points in KV space depending on its provenance, which refers to the existing KV entries within the caching window. When this provenance involves unseen context beyond the training horizon, the resulting KV entry falls out-of-distribution, a phenomenon termed the KV-provenance problem.
As illustrated in the framework diagram, generating a chunk from an out-of-distribution provenance causes subsequent generations to inherit this state, leading to progressive drifting. To resolve this, the authors introduce Self-Caching, a rule that restricts every provenance to configurations executed during training. For a conditioning window of length L, the first ℓ entries are cached by attending only to themselves, denoted as E(xj;∅). The remaining L−ℓ entries are cached autoregressively on top of the last self-cached entry. This ensures that no KV entry is ever written under an unseen context, keeping the provenance strictly in-distribution throughout extrapolation. The caching rule is formalized as:
κj={E(xj;∅),E(xj;κi−L+ℓ−1:j−1),i−L≤j<i−L+ℓ,i−L+ℓ≤j≤i−1.Level 2: KV Conditioning and Exact Rolling Window
While Self-Caching ensures stored KV entries are in-distribution, the authors note that how these entries are read must also match training. They address KV conditioning by treating the first chunk as a persistent condition and enabling an exact rolling window.
During training, every generation is conditioned on the first chunk KV entry, κ0. Because the causal 3D-VAE encodes the initial latent frame from a single pixel frame, κ0 is distributionally unique and irreplaceable. Evicting it during extrapolation would itself be out-of-distribution. Therefore, ID-Forcing retains κ0 at slot 0 as a persistent sink, differing from standard attention sinks that retain multiple early entries for empirical stability.
As shown in the figure below, the conditioning window is constrained to a length L<N, holding κ0 at slot 0 and the L−1 most recent KV entries in the remaining slots. This strictly bounds the total window size within the training budget. After generating each chunk, the oldest non-sink KV entry is evicted, and the remaining entries are re-rotated by R−f, while κ0 remains fixed. The conditioning rule is expressed as:
{κ0}∪{κi−L+1,…,κi−1}←{κ0}∪R−f[{κi−L+1,…,κi−1}].Because Self-Caching ensures that every cached KV entry possesses no external provenance to lose during eviction, this rolling window remains exactly in-distribution, matching the temporal distances observed during training.
Experiment
ID-Forcing is evaluated for minute-scale video generation beyond the training horizon on Self-Forcing and LongLive base models, using 120- and 240-second videos and comparing against Deep Forcing, ∞-RoPE, and MemRoPE with automated drift metrics and a user study. The main results show that ID-Forcing reduces color and motion drift while preserving subject consistency, visual quality, and dynamic motion, and qualitative comparisons confirm it avoids the collapse and artifacts seen in baselines. Ablations validate the contributions of self-caching, retaining the first chunk as a sink, and re-rotating the conditioning window, showing that these components together suppress drifting and that self-caching is especially important for stable long-horizon generation.
ID-Forcing shows the strongest long-duration stability among the compared methods, with the highest Dynamic Degree and clear gains on color and motion drift metrics in the reported settings. It also remains competitive on Aesthetic Quality, Imaging Quality, and Subject Consistency, while Motion Smoothness is lower primarily because that metric favors less dynamic outputs. Qualitative comparisons further indicate that ID-Forcing preserves subject identity, color fidelity, and motion over long videos. ID-Forcing achieves the highest Dynamic Degree in the reported settings, while Self-Forcing and LongLive baselines tend toward frame freezing at longer durations. It also shows substantial gains on color and motion drift metrics and remains competitive on quality and subject consistency, with Motion Smoothness as the only metric favoring less dynamic outputs.
In a pairwise user study, participants preferred ID-Forcing over all four Self-Forcing-based baselines across every evaluated dimension. The advantage was strongest against Self-Forcing, where preferences were near unanimous, including overall preference. Against MemRoPE, the closest dimension was dynamic degree, but ID-Forcing still received majority preference. ID-Forcing was preferred over every baseline on color consistency, background consistency, subject consistency, dynamic degree, temporal flicker, and overall preference. The strongest preference gap appeared against Self-Forcing, with near-unanimous ratings across most dimensions, while the smallest lead was on dynamic degree against MemRoPE.
Experiments evaluate ID-Forcing against Self-Forcing and LongLive baselines on long-duration video generation, using automated metrics, qualitative comparisons, and a pairwise user study. ID-Forcing achieves the strongest long-duration stability and highest Dynamic Degree, with clear gains on color and motion drift while remaining competitive on quality and subject consistency; Motion Smoothness is lower mainly because that metric favors less dynamic outputs. Qualitative results show preserved subject identity, color fidelity, and motion over long videos. In the user study, participants preferred ID-Forcing over all four Self-Forcing-based baselines across every evaluated dimension, including overall preference, with near-unanimous preference against Self-Forcing and the smallest lead on dynamic degree against MemRoPE.