Command Palette
Search for a command to run...
오디오 인지 스타일 참조를 통한 스타일 보존リップ 싱크
오디오 인지 스타일 참조를 통한 스타일 보존リップ 싱크
Weizhi Zhong Jichang Li Yinqi Cai Ming Li Feng Gao Liang Lin Guanbin Li
고품질 입맞춤 모델 MuseTalk 의 원클릭 배포
초록
멀티미디어 분야에서의 광범위한 응용 가능성으로 인해 오디오 기반 입모양 동기화(audio-driven lip sync) 기술은 최근 주목을 받고 있습니다. 같은 발화를 하더라도 개인마다 독특한 화법(style)을 보이며 이에 따라 입모양이 달라지기 때문에, 오디오 기반 입모양 동기화 작업은 상당한 과제를 안고 있습니다. 기존 방법론들은 종종 개인화된 화법 특성을 모델링하는 것을 우회하여, 일반적인 스타일에 기반한 최적화가 아닌 하위 최적(sub-optimal) 수준의 입모양 동기화를 초래했습니다. 최근의 입모양 동기화 기법들은 스타일 참조 영상에서 정보를 집계하여 임의의 오디오에 대한 입모양 동기화를 유도하고자 했으나, 스타일 집계 과정의 부정확성으로 인해 화법 특성을 충분히 유지하지 못하는 한계가 있었습니다.본 연구에서는 스타일 참조 영상에서 추출된 참조 오디오와 입력 오디오 간의 관계를 효과적으로 활용하여, 화법 특성을 보존하는 오디오 기반 입모양 동기화를 달성하기 위한 혁신적인 오디오 인식형 스타일 참조 방식을 제안합니다. 구체적으로, 먼저 스타일 참조 영상에서 cross-attention 레이어를 통해 집계된 스타일 정보를 결합하여, 입력 오디오에 상응하는 입모양을 정확하게 예측하는 데 능숙한 Transformer 기반 모델을 개발합니다. 이후, 예측된 입모양을 사실적인 화자 영상으로 더 잘 구현하기 위해, modulated convolutional 레이어를 통해 입모양을 통합하고 spatial cross-attention 레이어를 통해 참조 얼굴 이미지를 융합하는 조건부 latent diffusion 모델을 설계합니다. 광범위한 실험을 통해 제안된 접근 방식이 정밀한 입모양 동기화, 화법 특성 보존, 그리고 고충실도(high-fidelity)의 사실적인 화자 영상 생성에 효과적임을 입증하였습니다.
One-sentence Summary
Addressing the inaccurate style aggregation of prior methods, this work proposes an audio-aware style reference scheme that integrates a Transformer-based lip motion predictor enhanced by cross-attention layers for style aggregation and a conditional latent diffusion renderer fused via modulated convolutions and spatial cross-attention, with extensive experiments validating its ability to achieve precise lip synchronization, preserve individual speaking styles, and generate high-fidelity talking face videos.
Key Contributions
- This work proposes an audio-aware style reference scheme that models the relationship between input audio and reference audio to preserve individual speaking styles. A Transformer-based architecture predicts target lip motions by aggregating personalized style cues through cross-attention layers.
- A conditional latent diffusion model renders the predicted lip motions into realistic talking face videos. This renderer integrates motion signals through modulated convolutional layers and fuses reference facial images via spatial cross-attention mechanisms.
- Extensive experiments validate that the proposed framework achieves precise lip synchronization, effectively preserves individual speaking styles, and generates high-fidelity talking face videos. The results confirm the effectiveness of the integrated style aggregation and rendering pipeline.
Introduction
No source text was provided for analysis. Please share the abstract or body snippet so I can draft a concise research background summary that outlines the technical context, prior limitations, and the authors’ main contribution in a clear, technical yet readable format.