HyperAIHyperAI

Command Palette

Search for a command to run...

FRAMEMORROW : sélection de trames guidée par le futur à l'aide de jetons prospectifs pour la génération vidéo à long horizon

Bo Yin Xiaobin Hu Jiaqi Zhao Shuicheng Yan

Résumé

La génération vidéo à long horizon exige des modèles qu'ils exploitent efficacement un historique de génération de plus en plus long. À mesure que l'historique généré s'allonge, conserver tout le contenu antérieur devient de plus en plus coûteux et redondant, ce qui rend essentielle une sélection efficace de l'historique. Les approches existantes déterminent souvent la pertinence de l'historique à partir du contenu courant. Or, une information pertinente pour le présent n'est pas nécessairement utile pour la génération future, tandis qu'un historique apparemment moins pertinent peut devenir important par la suite. Notre idée clé est que l'information historique devrait être sélectionnée en fonction de sa pertinence pour les besoins futurs en information. Saisir ces besoins n'exige pas de générer entièrement le futur. Une représentation compacte de ce qui devient important ensuite suffit à guider la sélection historique. Sur la base de cette idée, nous proposons FRAMEMORROW, un sélecteur prospectif de trames qui prédit un petit ensemble de jetons prospectifs représentant les besoins futurs en information et les utilise pour identifier l'information pertinente dans l'historique. FRAMEMORROW sélectionne des trames historiques explicites plutôt que des états internes propres au modèle, ce qui permet une intégration enfichable dans divers générateurs, y compris des modèles à source fermée, avec un surcoût d'inférence minime. Nous évaluons FRAMEMORROW sur cinq bancs d'essai et 11 modèles génératifs couvrant la génération vidéo longue, la génération interactive et les modèles du monde conditionnés par l'action. Des expériences approfondies montrent des améliorations de la cohérence à longue portée, de la qualité visuelle et de l'alignement des actions dans divers contextes de génération. La page du projet est disponible à l'adresse https://yinbo0927.github.io/FrameMorrow/.

One-sentence Summary

Researchers from National University of Singapore and Harbin Institute of Technology (Shenzhen) propose FRAMEMORROW, a future-guided frame selector that predicts prospective tokens representing future information needs and selects explicit historical frames accordingly, enabling plug-and-play integration across diverse generators, including closed-source models, and improving long-range consistency, visual quality, and action alignment across five benchmarks and 11 generative models.

Key Contributions

  • Formulates historical frame selection for long-horizon video generation as conditioned on future information needs rather than on relevance to current content alone.
  • Introduces FrameMorrow, a plug-and-play prospective frame selector that predicts a small set of prospective tokens representing future information needs and selects explicit historical frames, enabling compatibility with diverse generators including closed-source models and little additional inference cost.
  • Evaluates FrameMorrow across five benchmarks and 11 generative models spanning long-video generation, interactive generation, and action-conditioned world models, showing improvements in long-range consistency, visual quality, and action alignment.

Introduction

Long-horizon video generation requires models to maintain coherent subjects, objects, and scenes over time, but autoregressive generators have a limited input window and cannot keep all earlier frames without high computation and redundant context. Existing history-selection methods typically retrieve past information by matching the current visual content, which captures what is relevant to the present but may miss earlier content that will become important later. The authors argue that historical relevance should be determined by future information needs, even though the future has not yet been generated. To address this, they propose FRAMEMORROW, which predicts a small set of prospective tokens as a compact proxy for future needs and uses them to select explicit historical frames. This makes the selector lightweight and plug-and-play across diverse generators, including closed-source models, and improves long-range consistency, visual quality, and action alignment in evaluations.

Method

The authors propose FRAMEMORROW, a generator-external prospective selector designed to identify and retrieve relevant historical frames for video generation. Instead of generating future content, the model predicts a compact set of prospective tokens that represent future information needs. These tokens evaluate the relevance of eligible historical frames and guide the selection process, ultimately augmenting a compatible frozen generator.

The framework operates by distinguishing the recent context from the eligible long-term history. At a given rollout step, the recent context informs the selection process without consuming the long-term memory budget, while the eligible history serves as the candidate pool for retrieval. A known rollout condition, such as a text prompt or action sequence, is also incorporated to guide the selection. The selector computes relevance scores for the eligible history and selects the top-K frames:

r^t=Sθ(Ht,Lt,ct+)∈RN,It=TopK(r^t,K)\hat{\mathbf{r}}_t = S_\theta(\mathcal{H}_t, \mathcal{L}_t, c_t^+) \in \mathbb{R}^N, \quad \mathcal{I}_t = \mathrm{TopK}(\hat{\mathbf{r}}_t, K)r^t​=Sθ​(Ht​,Lt​,ct+​)∈RN,It​=TopK(r^t​,K)

To represent future information needs, a causal Transformer autoregressively predicts a small set of prospective tokens. The eligible history, recent context, and rollout condition are first encoded using frozen visual and modality-specific encoders, followed by learned projections. The causal Transformer processes these encoded sequences and generates the prospective tokens:

qtm=Fψ(Ht,Lt,Ct+,qt<m),m=1,…,M\mathbf{q}_t^m = F_\psi(\mathbf{H}_t, \mathbf{L}_t, \mathbf{C}_t^+, \mathbf{q}_t^{<m}), \quad m = 1, \dots, Mqtm​=Fψ​(Ht​,Lt​,Ct+​,qt<m​),m=1,…,M

Each token can condition on its predecessors, allowing the predictions to capture complex dependencies. These prospective tokens then serve as attention queries over the eligible historical frames. The model computes scaled dot-product attention logits between the projected tokens and the encoded historical frames, aggregating them using a smooth maximum function to produce a final relevance score for each historical frame.

During training, the authors employ a future-grounded ranking distillation strategy to supervise the prospective tokens. A frozen visual teacher, specifically a DINOv2 encoder, evaluates the correspondence between candidate historical frames and the realized future continuation. By computing the cosine similarity between the encoded historical and future frames, the teacher derives a visual relevance proxy. These similarities are aggregated using a smooth maximum to favor frames that correspond to at least a portion of the future content. Both the teacher and student scores are normalized into ranking distributions. The selector is optimized using a combination of listwise and pairwise ranking losses:

Lsel=Llist+λLpair\mathcal{L}_{\mathrm{sel}} = \mathcal{L}_{\mathrm{list}} + \lambda \mathcal{L}_{\mathrm{pair}}Lsel​=Llist​+λLpair​

The listwise term aligns the full probability distribution of the student with the teacher, while the pairwise term preserves the relative orderings of frames separated by a specific margin in the teacher ranking. This distillation process ensures that the prospective tokens learn to identify historically relevant frames without requiring explicit future generation during inference.

At inference time, the continuation-based target construction is entirely removed. The selector encodes the observed frames and predicts prospective tokens using only the available history, recent context, and rollout condition. The selected explicit historical frames are then integrated into the generation pipeline through a plug-and-play mechanism. Each frozen backbone processes these selected frames through its own native conditioning mechanism, such as a reference-frame interface or memory adapter. This design allows the selector to dynamically decide which historical evidence to provide, while the generator retains its original architecture and processing methods, requiring no additional conditioning modules to be trained.

Experiment

The experiments evaluate FRAMEMORROW across five settings: long video generation, single-shot and multi-shot interactive generation, closed-source generators, and action-conditioned world models. Across open-source backbones, it consistently improves generation consistency and prompt alignment, with larger gains in later segments, and it remains effective through reference-conditioning interfaces for closed-source models. In world-model benchmarks, it strengthens visual consistency, anti-flicker, and action alignment while preserving motion quality. Ablations show that future-derived ranking targets, selecting relevant historical frames rather than simply increasing frame count, and four autoregressive prospective tokens provide most of the benefit, with only modest inference overhead.

FRAMEMORROW achieves the best average rank among long-context methods on all three evaluated backbones for 60-second generation. On Self-Forcing, it delivers notable gains in imaging quality and dynamic degree, while LongLive-RAG keeps higher background and motion scores. The consistent average-rank advantage supports the effectiveness of frame selection across generators, with trade-offs in individual metrics. FRAMEMORROW achieves the best average rank on the Self-Forcing, LongLive 1.0, and Causal Forcing backbones. On Self-Forcing, FRAMEMORROW improves imaging quality from 62.22 to 68.47 and dynamic degree from 51.72 to 64.56. LongLive-RAG retains higher background and motion scores on Self-Forcing despite FRAMEMORROW's better average rank.

In single-shot interactive video generation over 60 seconds, FRAMEMORROW improves overall quality, consistency, and aesthetic scores over the reported native backbones. Segment-wise CLIP score gains are larger in later 10-second intervals, particularly for Self-Forcing, indicating stronger prompt alignment at later stages. For single-shot generation, FRAMEMORROW improves overall quality and consistency over LongLive 1.0 and Self-Forcing baselines. With Self-Forcing, consistency rises from 84.95 to 89.30, while segment-wise CLIP-score gains are smallest in the first 10 seconds and largest in the final 10 seconds. Aesthetic scores also improve modestly for the reported backbone pairs, and later 10-second windows show larger prompt-alignment gains.

In closed-source generation tests across Seedance 2.0 and Kling O3, FRAMEMORROW improves quality, consistency, and aesthetic scores over both the original generators and uniform historical-frame selection when using the same reference budget. Uniform selection slightly degrades all reported metrics relative to the native baselines. The results indicate the approach can work through public reference-conditioning interfaces without access to generator internals. FRAMEMORROW outperforms uniform selection on all reported metrics for both closed-source generators. Uniform historical-frame selection lowers quality, consistency, and aesthetic scores compared with the original closed-source baselines. FRAMEMORROW achieves consistency gains over the native generators on Seedance 2.0 and Kling O3.

FRAMEMORROW consistently improves subject consistency, background consistency, imaging quality, anti-flicker, and action alignment across Matrix-Game 3.0, WorldMem, and YuMe 1.5. Motion quality remains largely stable, with only small changes in either direction. The gains occur within each base-model pair, indicating that the method benefits diverse interactive world models. FRAMEMORROW raises action alignment for all three backbones, with the largest improvement on YuMe 1.5. Subject and background consistency, imaging quality, and anti-flicker improve over the corresponding base model in every evaluated pair. Motion scores stay nearly unchanged, suggesting the method preserves motion quality while improving consistency and control alignment.

FRAMEMORROW outperforms all five historical frame selection strategies in both single-shot and multi-shot generation. Prompt matching is the strongest baseline, followed by direct scoring and context matching, while simple recency and uniform sampling lag behind. The performance gaps are more pronounced in single-shot generation than in multi-shot generation. FRAMEMORROW achieves the highest consistency scores in both single-shot and multi-shot settings. Prompt matching is the best-performing baseline, with direct scoring and context matching also competitive. Simple recency and uniform sampling trail the learned matching and scoring strategies. All strategies improve in multi-shot generation compared with single-shot generation. FRAMEMORROW's advantage over the strongest baseline is larger in single-shot generation than in multi-shot generation.

Across experiments, FRAMEMORROW was evaluated on long-context generation with multiple backbones, single-shot interactive video, closed-source APIs through public reference-conditioning interfaces, diverse interactive world models, and comparisons against several historical-frame selection baselines. It consistently achieves the best average rank or highest consistency and improves subject consistency, background consistency, imaging quality, aesthetic scores, and action alignment over native baselines, with motion quality remaining stable and occasional trade-offs in background or motion scores. Gains in prompt alignment are stronger in later 10-second intervals, and the approach outperforms uniform, recency, and learned scoring baselines, particularly in single-shot generation. Overall, the results indicate that FRAMEMORROW's frame selection benefits diverse generators and closed-source settings without requiring access to generator internals.


Créer de l'IA avec l'IA

De l'idée au lancement — accélérez votre développement IA avec le co-codage IA gratuit, un environnement prêt à l'emploi et le meilleur prix pour les GPU.

Codage assisté par IA
GPU prêts à l’emploi
Tarifs les plus avantageux

HyperAI Newsletters

Abonnez-vous à nos dernières mises à jour
Nous vous enverrons les dernières mises à jour de la semaine dans votre boîte de réception à neuf heures chaque lundi matin
Propulsé par MailChimp