RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911469143064576 |
|---|---|
| author | Du, Fangyu Li, Taiqing Qiao, Qian Yu, Tan Zhang, Ziwei Zhen, Dingcheng Jia, Xu Yang, Yang Yin, Shunshun Liu, Siyuan |
| author_facet | Du, Fangyu Li, Taiqing Qiao, Qian Yu, Tan Zhang, Ziwei Zhen, Dingcheng Jia, Xu Yang, Yang Yin, Shunshun Liu, Siyuan |
| contents | Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve high-quality results by leveraging high-dimensional intermediate representations and explicitly modeling motion dynamics, their computational complexity renders them unsuitable for real-time deployment. Real-time inference imposes stringent latency and memory constraints, often necessitating the use of highly compressed latent representations. However, operating in such compact spaces hinders the preservation of fine-grained spatiotemporal details, thereby complicating audio-visual synchronization RAP (Real-time Audio-driven Portrait animation), a unified framework for generating high-quality talking portraits under real-time constraints. Specifically, RAP introduces a hybrid attention mechanism for fine-grained audio control, and a static-dynamic training-inference paradigm that avoids explicit motion supervision. Through these techniques, RAP achieves precise audio-driven control, mitigates long-term temporal drift, and maintains high visual fidelity. Extensive experiments demonstrate that RAP achieves state-of-the-art performance while operating under real-time constraints. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_05115 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer Du, Fangyu Li, Taiqing Qiao, Qian Yu, Tan Zhang, Ziwei Zhen, Dingcheng Jia, Xu Yang, Yang Yin, Shunshun Liu, Siyuan Graphics Computer Vision and Pattern Recognition Sound Audio and Speech Processing Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve high-quality results by leveraging high-dimensional intermediate representations and explicitly modeling motion dynamics, their computational complexity renders them unsuitable for real-time deployment. Real-time inference imposes stringent latency and memory constraints, often necessitating the use of highly compressed latent representations. However, operating in such compact spaces hinders the preservation of fine-grained spatiotemporal details, thereby complicating audio-visual synchronization RAP (Real-time Audio-driven Portrait animation), a unified framework for generating high-quality talking portraits under real-time constraints. Specifically, RAP introduces a hybrid attention mechanism for fine-grained audio control, and a static-dynamic training-inference paradigm that avoids explicit motion supervision. Through these techniques, RAP achieves precise audio-driven control, mitigates long-term temporal drift, and maintains high visual fidelity. Extensive experiments demonstrate that RAP achieves state-of-the-art performance while operating under real-time constraints. |
| title | RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer |
| topic | Graphics Computer Vision and Pattern Recognition Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2508.05115 |