HyperAIHyperAI

Command Palette

Search for a command to run...

In-Distribution Forcing für die Generierung langer Videos zur Testzeit

Jeongwoo Shin Youngyoon Choi Sangwoo Jo Hyunmog Kim Sungjoon Choi Joonseok Lee Jaewoong Choi Jaemoo Choi

Zusammenfassung

Moderne autoregressive (AR) Videodiffusionsmodelle beherrschen die Generierung kurzer Videos hervorragend, doch die Generierung langer Videos bleibt aufgrund von Drift schwierig, wobei sich Farben und Texturen verschieben und die Bewegungsdynamik nachlässt. Bestehende Arbeiten stützen sich überwiegend auf KV-Konditionierung, die zwischengespeicherte Key-Value-Einträge (KV) auswählt oder modifiziert, um Drift zu verringern. Wir beobachten jedoch, dass KV-Konditionierung allein nicht ausreicht, da sie annimmt, dass zwischengespeicherte KV-Einträge in der Verteilung bleiben. Diese Annahme versagt jenseits des Trainingshorizonts: Nichts beschränkt die Konstruktion von KV-Einträgen während des Rollouts, wodurch das KV-Provenienz-Problem entsteht, bei dem die zwischengespeicherten Einträge selbst Out-of-Distribution (OOD) werden. Um dies zu beheben, schlagen wir In-Distribution Forcing (ID-Forcing) vor, ein Testzeit-Framework, das sowohl KV-Caching als auch KV-Konditionierung mit den Trainingskonfigurationen in Einklang bringt. Sein zentraler Mechanismus, Self-Caching, verhindert OOD-KV-Einträge an ihrer Quelle. Jeder Chunk wird zwischengespeichert, ohne dass dabei Aufmerksamkeit auf frühere KV-Einträge gelegt wird; dadurch bleibt das gleitende Fenster exakt in der Verteilung. Infolgedessen erweitert ID-Forcing Kurzzeitmodelle nahtlos auf die Videogenerierung im Minutenbereich. Umfangreiche Evaluierungen zeigen, dass unsere Methode auf Standard-Benchmarks für Videogenerierung wettbewerbsfähig bleibt und frühere Arbeiten bei der Eindämmung von Drift deutlich übertrifft, was sowohl durch unsere Driftmetriken als auch durch eine Nutzerstudie bestätigt wird.

One-sentence Summary

Researchers from Seoul National University, Korea University, Sungkyunkwan University, and Georgia Institute of Technology propose In-Distribution Forcing (ID-Forcing), a test-time framework that aligns both KV caching and KV conditioning with training configurations and uses a selfcaching mechanism to prevent out-of-distribution key-value entries at their source, outperforming prior KV conditioning methods in mitigating drifting and extending short-horizon video diffusion models to minute-scale generation.

Key Contributions

  • The paper identifies the KV-provenance problem in autoregressive video diffusion: beyond the training horizon, cached key-value entries from rolling-window generation become out-of-distribution, so KV conditioning alone cannot prevent drifting.
  • It introduces a training-free test-time framework, In-Distribution Forcing (ID-Forcing), whose self-caching mechanism caches each chunk without attending to prior KV entries; at the conditioning level it bounds the window length to L ≤ N−1 and permanently fixes the first chunk κ0.
  • Evaluations across durations, backbones, and metrics show that ID-Forcing remains competitive on a standard video generation benchmark and substantially outperforms prior test-time approaches in mitigating drifting, as measured by drift metrics and a user study, including minute-scale generation.

Introduction

Autoregressive video diffusion enables long-horizon synthesis by generating video chunks sequentially and conditioning each chunk on key-value cache entries from earlier chunks, which matters for streaming, world models, and long-form storytelling. A key limitation is drifting: beyond the finite training horizon, colors and textures shift and motion decays because test-time rollouts enter conditioning configurations not seen during training. Prior methods mostly adjust KV conditioning, such as attention sinks or temporal position embeddings, but they assume the cached KV entries themselves remain in-distribution. The authors show this assumption can fail because each KV entry is built from its own attended context, called its provenance; beyond the training horizon, caching itself uses unseen provenance. Their main contribution is In-Distribution Forcing, which uses self-caching and a bounded rolling window so that both KV caching and KV conditioning stay within training-time configurations, allowing a short video model to generate minute-scale video with reduced drift.

Method

The authors propose In-Distribution Forcing (ID-Forcing), a training-free framework designed to keep test-time extrapolation strictly within the training distribution. Unlike prior works that treat drifting as a consequence of discarding informative past context and only modify KV conditioning, ID-Forcing views drifting as a deviation from the training distribution that occurs in both KV caching and KV conditioning. Therefore, the framework introduces two-level control over KV operations to prevent out-of-distribution states from accumulating.

Level 1: KV Caching and the KV-Provenance Problem

The authors identify that the same clean chunk maps to different points in KV space depending on its provenance, which refers to the existing KV entries within the caching window. When this provenance involves unseen context beyond the training horizon, the resulting KV entry falls out-of-distribution, a phenomenon termed the KV-provenance problem.

As illustrated in the framework diagram, generating a chunk from an out-of-distribution provenance causes subsequent generations to inherit this state, leading to progressive drifting. To resolve this, the authors introduce Self-Caching, a rule that restricts every provenance to configurations executed during training. For a conditioning window of length LLL, the first ℓ\ellℓ entries are cached by attending only to themselves, denoted as E(xj;∅)E(x_j; \varnothing)E(xj​;∅). The remaining L−ℓL - \ellL−ℓ entries are cached autoregressively on top of the last self-cached entry. This ensures that no KV entry is ever written under an unseen context, keeping the provenance strictly in-distribution throughout extrapolation. The caching rule is formalized as:

κj={E(xj;∅),i−L≤j<i−L+ℓ,E(xj;κi−L+ℓ−1:j−1),i−L+ℓ≤j≤i−1.\kappa_{j} = \left\{ \begin{array}{l l} E (x_{j}; \varnothing), & i - L \leq j < i - L + \ell, \\ E (x_{j}; \kappa_{i - L + \ell - 1: j - 1}), & i - L + \ell \leq j \leq i - 1. \end{array} \right.κj​={E(xj​;∅),E(xj​;κi−L+ℓ−1:j−1​),​i−L≤j<i−L+ℓ,i−L+ℓ≤j≤i−1.​

Level 2: KV Conditioning and Exact Rolling Window

While Self-Caching ensures stored KV entries are in-distribution, the authors note that how these entries are read must also match training. They address KV conditioning by treating the first chunk as a persistent condition and enabling an exact rolling window.

During training, every generation is conditioned on the first chunk KV entry, κ0\kappa_0κ0​. Because the causal 3D-VAE encodes the initial latent frame from a single pixel frame, κ0\kappa_0κ0​ is distributionally unique and irreplaceable. Evicting it during extrapolation would itself be out-of-distribution. Therefore, ID-Forcing retains κ0\kappa_0κ0​ at slot 0 as a persistent sink, differing from standard attention sinks that retain multiple early entries for empirical stability.

As shown in the figure below, the conditioning window is constrained to a length L<NL < NL<N, holding κ0\kappa_0κ0​ at slot 0 and the L−1L - 1L−1 most recent KV entries in the remaining slots. This strictly bounds the total window size within the training budget. After generating each chunk, the oldest non-sink KV entry is evicted, and the remaining entries are re-rotated by R−fR_{-f}R−f​, while κ0\kappa_0κ0​ remains fixed. The conditioning rule is expressed as:

{κ0}∪{κi−L+1,…,κi−1}←{κ0}∪R−f[{κi−L+1,…,κi−1}].\{\kappa_{0}\} \cup \{\kappa_{i - L + 1}, \dots, \kappa_{i - 1}\} \leftarrow \{\kappa_{0}\} \cup R_{-f} \big[ \{\kappa_{i - L + 1}, \dots, \kappa_{i - 1}\} \big].{κ0​}∪{κi−L+1​,…,κi−1​}←{κ0​}∪R−f​[{κi−L+1​,…,κi−1​}].

Because Self-Caching ensures that every cached KV entry possesses no external provenance to lose during eviction, this rolling window remains exactly in-distribution, matching the temporal distances observed during training.

Experiment

ID-Forcing is evaluated for minute-scale video generation beyond the training horizon on Self-Forcing and LongLive base models, using 120- and 240-second videos and comparing against Deep Forcing, ∞-RoPE, and MemRoPE with automated drift metrics and a user study. The main results show that ID-Forcing reduces color and motion drift while preserving subject consistency, visual quality, and dynamic motion, and qualitative comparisons confirm it avoids the collapse and artifacts seen in baselines. Ablations validate the contributions of self-caching, retaining the first chunk as a sink, and re-rotating the conditioning window, showing that these components together suppress drifting and that self-caching is especially important for stable long-horizon generation.

ID-Forcing shows the strongest long-duration stability among the compared methods, with the highest Dynamic Degree and clear gains on color and motion drift metrics in the reported settings. It also remains competitive on Aesthetic Quality, Imaging Quality, and Subject Consistency, while Motion Smoothness is lower primarily because that metric favors less dynamic outputs. Qualitative comparisons further indicate that ID-Forcing preserves subject identity, color fidelity, and motion over long videos. ID-Forcing achieves the highest Dynamic Degree in the reported settings, while Self-Forcing and LongLive baselines tend toward frame freezing at longer durations. It also shows substantial gains on color and motion drift metrics and remains competitive on quality and subject consistency, with Motion Smoothness as the only metric favoring less dynamic outputs.

In a pairwise user study, participants preferred ID-Forcing over all four Self-Forcing-based baselines across every evaluated dimension. The advantage was strongest against Self-Forcing, where preferences were near unanimous, including overall preference. Against MemRoPE, the closest dimension was dynamic degree, but ID-Forcing still received majority preference. ID-Forcing was preferred over every baseline on color consistency, background consistency, subject consistency, dynamic degree, temporal flicker, and overall preference. The strongest preference gap appeared against Self-Forcing, with near-unanimous ratings across most dimensions, while the smallest lead was on dynamic degree against MemRoPE.

Experiments evaluate ID-Forcing against Self-Forcing and LongLive baselines on long-duration video generation, using automated metrics, qualitative comparisons, and a pairwise user study. ID-Forcing achieves the strongest long-duration stability and highest Dynamic Degree, with clear gains on color and motion drift while remaining competitive on quality and subject consistency; Motion Smoothness is lower mainly because that metric favors less dynamic outputs. Qualitative results show preserved subject identity, color fidelity, and motion over long videos. In the user study, participants preferred ID-Forcing over all four Self-Forcing-based baselines across every evaluated dimension, including overall preference, with near-unanimous preference against Self-Forcing and the smallest lead on dynamic degree against MemRoPE.


KI mit KI entwickeln

Von der Idee bis zum Launch – beschleunigen Sie Ihre KI-Entwicklung mit kostenlosem KI-Co-Coding, sofort einsatzbereiter Umgebung und bestem GPU-Preis.

KI-gestütztes kollaboratives Programmieren
Sofort einsatzbereite GPUs
Die besten Preise

HyperAI Newsletters

Abonnieren Sie unsere neuesten Updates
Wir werden die neuesten Updates der Woche in Ihren Posteingang liefern um neun Uhr jeden Montagmorgen
Unterstützt von MailChimp