InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Shaoshu, Kong, Zhe, Gao, Feng, Cheng, Meng, Liu, Xiangyu, Zhang, Yong, Kang, Zhuoliang, Luo, Wenhan, Cai, Xunliang, He, Ran, Wei, Xiaoming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909743690285056
author Yang, Shaoshu
Kong, Zhe
Gao, Feng
Cheng, Meng
Liu, Xiangyu
Zhang, Yong
Kang, Zhuoliang
Luo, Wenhan
Cai, Xunliang
He, Ran
Wei, Xiaoming
author_facet Yang, Shaoshu
Kong, Zhe
Gao, Feng
Cheng, Meng
Liu, Xiangyu
Zhang, Yong
Kang, Zhuoliang
Luo, Wenhan
Cai, Xunliang
He, Ran
Wei, Xiaoming
contents Recent breakthroughs in video AIGC have ushered in a transformative era for audio-driven human animation. However, conventional video dubbing techniques remain constrained to mouth region editing, resulting in discordant facial expressions and body gestures that compromise viewer immersion. To overcome this limitation, we introduce sparse-frame video dubbing, a novel paradigm that strategically preserves reference keyframes to maintain identity, iconic gestures, and camera trajectories while enabling holistic, audio-synchronized full-body motion editing. Through critical analysis, we identify why naive image-to-video models fail in this task, particularly their inability to achieve adaptive conditioning. Addressing this, we propose InfiniteTalk, a streaming audio-driven generator designed for infinite-length long sequence dubbing. This architecture leverages temporal context frames for seamless inter-chunk transitions and incorporates a simple yet effective sampling strategy that optimizes control strength via fine-grained reference frame positioning. Comprehensive evaluations on HDTF, CelebV-HQ, and EMTD datasets demonstrate state-of-the-art performance. Quantitative metrics confirm superior visual realism, emotional coherence, and full-body motion synchronization.
format Preprint
id arxiv_https___arxiv_org_abs_2508_14033
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
Yang, Shaoshu
Kong, Zhe
Gao, Feng
Cheng, Meng
Liu, Xiangyu
Zhang, Yong
Kang, Zhuoliang
Luo, Wenhan
Cai, Xunliang
He, Ran
Wei, Xiaoming
Computer Vision and Pattern Recognition
Recent breakthroughs in video AIGC have ushered in a transformative era for audio-driven human animation. However, conventional video dubbing techniques remain constrained to mouth region editing, resulting in discordant facial expressions and body gestures that compromise viewer immersion. To overcome this limitation, we introduce sparse-frame video dubbing, a novel paradigm that strategically preserves reference keyframes to maintain identity, iconic gestures, and camera trajectories while enabling holistic, audio-synchronized full-body motion editing. Through critical analysis, we identify why naive image-to-video models fail in this task, particularly their inability to achieve adaptive conditioning. Addressing this, we propose InfiniteTalk, a streaming audio-driven generator designed for infinite-length long sequence dubbing. This architecture leverages temporal context frames for seamless inter-chunk transitions and incorporates a simple yet effective sampling strategy that optimizes control strength via fine-grained reference frame positioning. Comprehensive evaluations on HDTF, CelebV-HQ, and EMTD datasets demonstrate state-of-the-art performance. Quantitative metrics confirm superior visual realism, emotional coherence, and full-body motion synchronization.
title InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.14033