RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Fangyu, Li, Taiqing, Qiao, Qian, Yu, Tan, Zhang, Ziwei, Zhen, Dingcheng, Jia, Xu, Yang, Yang, Yin, Shunshun, Liu, Siyuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911469143064576
author Du, Fangyu
Li, Taiqing
Qiao, Qian
Yu, Tan
Zhang, Ziwei
Zhen, Dingcheng
Jia, Xu
Yang, Yang
Yin, Shunshun
Liu, Siyuan
author_facet Du, Fangyu
Li, Taiqing
Qiao, Qian
Yu, Tan
Zhang, Ziwei
Zhen, Dingcheng
Jia, Xu
Yang, Yang
Yin, Shunshun
Liu, Siyuan
contents Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve high-quality results by leveraging high-dimensional intermediate representations and explicitly modeling motion dynamics, their computational complexity renders them unsuitable for real-time deployment. Real-time inference imposes stringent latency and memory constraints, often necessitating the use of highly compressed latent representations. However, operating in such compact spaces hinders the preservation of fine-grained spatiotemporal details, thereby complicating audio-visual synchronization RAP (Real-time Audio-driven Portrait animation), a unified framework for generating high-quality talking portraits under real-time constraints. Specifically, RAP introduces a hybrid attention mechanism for fine-grained audio control, and a static-dynamic training-inference paradigm that avoids explicit motion supervision. Through these techniques, RAP achieves precise audio-driven control, mitigates long-term temporal drift, and maintains high visual fidelity. Extensive experiments demonstrate that RAP achieves state-of-the-art performance while operating under real-time constraints.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05115
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer
Du, Fangyu
Li, Taiqing
Qiao, Qian
Yu, Tan
Zhang, Ziwei
Zhen, Dingcheng
Jia, Xu
Yang, Yang
Yin, Shunshun
Liu, Siyuan
Graphics
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve high-quality results by leveraging high-dimensional intermediate representations and explicitly modeling motion dynamics, their computational complexity renders them unsuitable for real-time deployment. Real-time inference imposes stringent latency and memory constraints, often necessitating the use of highly compressed latent representations. However, operating in such compact spaces hinders the preservation of fine-grained spatiotemporal details, thereby complicating audio-visual synchronization RAP (Real-time Audio-driven Portrait animation), a unified framework for generating high-quality talking portraits under real-time constraints. Specifically, RAP introduces a hybrid attention mechanism for fine-grained audio control, and a static-dynamic training-inference paradigm that avoids explicit motion supervision. Through these techniques, RAP achieves precise audio-driven control, mitigates long-term temporal drift, and maintains high visual fidelity. Extensive experiments demonstrate that RAP achieves state-of-the-art performance while operating under real-time constraints.
title RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer
topic Graphics
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2508.05115