HyperAIHyperAI

Command Palette

Search for a command to run...

التوجيه التكيفي للمكافأة: تحسين ديناميكي متعدد المكافآت لنماذج الانتشار المشتركة للصوت والفيديو عبر التعلم المعزز للعملية الأمامية

Songlin Yang Xiaotong Zhao Jiacheng Zhang Zhe Wang Toyota Li Eric Liu Alan Zhao Anyi Rao

الملخص

يُقدِّم التعلم المعزز الموجَّه بمكافآت متعددة (أي RL) طريقًا واعدًا لتحسين نماذج الانتشار المشتركة للصوت والفيديو عبر عدة أهداف، تشمل جودة كل نمط على حدة، والمواءمة الدلالية عبر الوسائط، والمزامنة الزمنية. غير أن فعاليته تعتمد على كميتين تتغيران أثناء التدريب: أين ينبغي أن تعمل التحديثات المدفوعة بالمكافأة، وكيف ينبغي تنسيق المكافآت المتنافسة. تميل الطرق الحالية إلى الاعتماد على توجيه ثابت وأوزان مكافآت ثابتة، مما لا يمكنها من تتبع دوال النموذج المتطورة. ولمعالجة هذه القيود، نقترح «التوجيه التكيفي للمكافأة» لتكييف مواضع التحديث وتنسيق المكافآت معًا أثناء التعلم المعزز للعملية الأمامية (أي DiffusionNFT) لنماذج الانتشار المشتركة للصوت والفيديو. تتكون طريقتنا من عنصرين. (1) توجيه يسترشد بالتأثير عبر الوسائط (تحديد مواضع التحديث): نستخدم استجابات الانتباه المتقاطع ثنائي الاتجاه بوصفها مؤشرًا وسيطًا كفؤًا للتأثير المتطور عبر الوسائط، مع إعادة ترجيح خسائر واعية بالرموز ديناميكيًا وتوسيع نطاق التدرجات عبر الطبقات عابرة الوسائط دون تدخلات إضافية في النموذج. (2) إعادة ترجيح حافظة للتفضيلات واعية بالأنماط (تنسيق المكافآت): نحافظ على الأوزان المحددة مسبقًا بوصفها معلومات مسبقة للتفضيلات، ونستخدم تفاعلات تدرج المكافأة الخاصة بكل فرع بوصفها تصحيحات متبقية بعد فترة الإحماء. وهذا يحل النزاعات المتطورة دون السماح للمكافآت المهيمنة بقمع الأهداف الضعيفة لكن الأساسية. تُظهر تجارب موسعة تحسينات متسقة في جودة الأنماط، والاتساق الدلالي، ومزامنة الصوت والفيديو مقارنةً بخطوط أساس قوية للتعلم المعزز. كما تدعم دراسات الاستئصال وتحليلات الآلية الفوائد التكاملية للتوجيه التكيفي للتحديث وتنسيق المكافآت.

One-sentence Summary

Researchers from Tencent Video and The University of Hong Kong propose Adaptive Reward Routing, a method that jointly adapts update locations and reward coordination during forward-process RL (DiffusionNFT) for joint audio-video diffusion models via cross-modal influence-guided routing and preference-preserving modality-aware reweighting, thereby improving modality quality, semantic consistency, and audio-video synchronization over strong RL baselines.

Key Contributions

  • Adaptive Reward Routing is a dynamic credit-assignment method for multi-reward reinforcement learning in joint audio-video diffusion models that adapts both where reward-based updates are applied and how competing reward signals are coordinated during forward-process RL.
  • The method introduces cross-modal influence-guided routing, which uses bidirectional cross-attention responses to dynamically reweight token-aware losses and scale gradients across crossmodal layers, and preference-preserving modality-aware reweighting, which keeps predefined reward weights as priors and adds branch-specific reward-gradient corrections after warm-up.
  • Experiments with dual-stream LTX and unified JavisDiT++ backbones show consistent improvements in modality quality, semantic consistency, and audio-video synchronization over strong RL baselines, while ablations and mechanism analyses validate the complementary benefit of adaptive update routing and reward coordination.

Introduction

The authors address post-training of joint audio-video diffusion models, where generated content must simultaneously satisfy modality-specific quality, cross-modal semantic alignment, and temporal synchronization. Reward-guided diffusion RL is appealing because multiple reward signals can express these requirements, but prior work relies on static update routing or fixed reward weights that become stale as the model evolves and reward conflicts shift. The authors propose Adaptive Reward Routing, which dynamically localizes reward-driven updates using bidirectional cross-attention responses and adapts reward coordination as residual corrections to user-defined preference weights. Experiments show consistent improvements in modality quality, semantic alignment, and audio-video synchronization across different backbones.

Method

The authors propose Adaptive Reward Routing, a forward-process reinforcement learning framework designed for the reward-guided post-training of joint audio-video diffusion models. This method addresses multi-modal and multi-reward optimization by dynamically adapting both reward priorities and routing locations.

The overall pipeline consists of five sequential stages: sampling, evaluation, reweighting, routing, and loss computation. During the sampling phase, the base model processes separate audio and video streams under a shared timestep. These streams exchange information through bidirectional cross-attention mechanisms, specifically audio-to-video and video-to-audio pathways, allowing the model to predict joint velocity fields. The generated samples are then evaluated using a suite of video, audio, and cross-modal rewards.

To effectively combine these diverse rewards, the framework employs a Preference-Preserving Modality-Aware Reweighting module. Instead of relying solely on predefined static weights or purely gradient-based coefficients, the authors combine both approaches. Each reward is probed through its corresponding modality branch. After an initial warm-up phase, a smoothed conflict-aware coefficient provides a residual correction to the prior weights. This ensures that the prior preferences establish a nonzero floor, while the residual term dynamically adapts to current gradient conflicts. The reweighted rewards are aggregated to compute a separate advantage AmA_mAm​ for each modality branch m∈{v,a}m \in \{v, a\}m∈{v,a}.

Following reweighting, the framework applies Cross-Modal Influence-Guided Routing to determine where the optimization updates should act. This routing operates at two distinct levels. At the token level, the model measures the directional influence of cross-modal attention by evaluating the pre-gate response norms. These responses are averaged over intermediate-to-late denoising timesteps and normalized to produce positive loss weights λm,i\lambda_{m,i}λm,i​ for individual tokens. At the layer level, the same responses are averaged across tokens to compute a layer score. This score is converted into a soft detachment coefficient that scales the backward gradient of the cross-attention keys and values. Consequently, strongly influential layers retain more gradient flow, while weakly coupled layers are increasingly detached, without altering the forward pass values.

Finally, the training objective integrates these adaptive mechanisms. The token routing weights are applied to a negative-aware loss function:

ℓm,i=rm∥vθ,m,i+−um,i∥22wm++ε+(1−rm)∥vθ,m,i−−um,i∥22wm−+ε\ell_{m,i} = r_m \frac{\| v_{\theta,m,i}^+ - u_{m,i} \|_2^2}{w_m^+ + \varepsilon} + (1 - r_m) \frac{\| v_{\theta,m,i}^- - u_{m,i} \|_2^2}{w_m^- + \varepsilon}ℓm,i​=rm​wm+​+ε∥vθ,m,i+​−um,i​∥22​​+(1−rm​)wm−​+ε∥vθ,m,i−​−um,i​∥22​​

where rmr_mrm​ is the optimality probability derived from the branch advantage, and wm±w_m^\pmwm±​ represents the detached mean absolute residual of the corresponding policy. The final loss combines the policy losses from both the audio and video branches and regularizes them toward a fixed reference policy using a Kullback-Leibler divergence term. This comprehensive design ensures that the optimization process respects user preferences while dynamically resolving multi-modal conflicts.

Experiment

The experiments evaluate adaptive reward routing on LTX-2 and LTX-2.3 audio-video diffusion backbones using a VGGSound-derived training set and the JavisBench benchmark. The main results show that the proposed method improves audio-video quality, text consistency, cross-modal consistency, and synchronization more effectively than fixed reward coordination, conflict-aware aggregation, and static multimodal routing baselines. Ablations confirm that token-level and layer-level routing are complementary, and that preference-preserving reward coordination addresses distinct failure modes. Additional validation demonstrates that the cross-modal influence proxy reliably identifies functionally important layers and tokens, while a frozen routing map becomes stale during training.

Under the reported LTX-2 comparisons, the proposed method improves substantially over the base model and existing baselines across visual quality, audio quality, text alignment, audio-visual consistency, and synchronization. GDPO and MARBLE offer partial gains but can sacrifice synchronization or text alignment. The complete model achieves the strongest overall balance, with leading scores across the reported metrics and the lowest desynchronization score. The proposed method attains the highest visual quality, audio quality, text consistency, audio-visual consistency, and JavisScore among the compared LTX-2 configurations, while also achieving the lowest desynchronization score. GDPO alone improves visual quality relative to the base model but increases desynchronization, showing that reward optimization without the full approach can sacrifice audio-visual timing. OmniNFT substantially improves audio-visual consistency and synchronization over earlier baselines, and the proposed method further improves these alignment metrics and overall quality.

In the LTX-2 backbone ablation, combining token weighting and layer scaling yields the strongest routing-only configuration, improving audiovisual quality, semantic consistency, and synchronization while reducing desynchronization. The isolated reward-coordination chain shows that branch-aware balancing, residual mixing, and warm-up progressively improve overall balance, with warm-up notably raising audio quality and lowering desync. Routing and weighting address complementary failure modes, and the complete model performs best overall. Token weighting and layer scaling are complementary: their combination produces the strongest routing-only results across quality, consistency, and synchronization metrics. Among isolated reward-coordination variants, adding warm-up to branch-aware balancing and residual mixing achieves the best audio quality and lowest desynchronization, though intermediate gains are not monotonic across every metric.

The experiments evaluate the proposed method against LTX-2 baselines and ablations, showing consistent improvements in visual quality, audio quality, text alignment, audio-visual consistency, and synchronization while achieving the lowest desynchronization score. Baseline methods such as GDPO and MARBLE offer partial gains but can sacrifice synchronization or text alignment, and OmniNFT improves consistency and synchronization but remains below the complete model. Ablations indicate that token weighting and layer scaling are complementary for routing, while branch-aware balancing, residual mixing, and warm-up progressively improve overall balance, with warm-up especially raising audio quality and reducing desynchronization.


بناء الذكاء الاصطناعي بالذكاء الاصطناعي

من الفكرة إلى الإطلاق — سرّع تطوير الذكاء الاصطناعي الخاص بك مع المساعدة البرمجية المجانية بالذكاء الاصطناعي، وبيئة جاهزة للاستخدام، وأفضل أسعار لوحدات معالجة الرسومات.

البرمجة التعاونية باستخدام الذكاء الاصطناعي
وحدات GPU جاهزة للعمل
أفضل الأسعار

HyperAI Newsletters

اشترك في آخر تحديثاتنا
سنرسل لك أحدث التحديثات الأسبوعية إلى بريدك الإلكتروني في الساعة التاسعة من صباح كل يوم اثنين
مدعوم بواسطة MailChimp