InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Shaoshu, Kong, Zhe, Gao, Feng, Cheng, Meng, Liu, Xiangyu, Zhang, Yong, Kang, Zhuoliang, Luo, Wenhan, Cai, Xunliang, He, Ran, Wei, Xiaoming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
by: Kong, Zhe, et al.
Published: (2025)
by: Kong, Zhe, et al.
Published: (2025)
DAM-VSR: Disentanglement of Appearance and Motion for Video Super-Resolution
by: Kong, Zhe, et al.
Published: (2025)
by: Kong, Zhe, et al.
Published: (2025)
Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory
by: Wu, Ruiqi, et al.
Published: (2026)
by: Wu, Ruiqi, et al.
Published: (2026)
ZeroSmooth: Training-free Diffuser Adaptation for High Frame Rate Video Generation
by: Yang, Shaoshu, et al.
Published: (2024)
by: Yang, Shaoshu, et al.
Published: (2024)
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
by: Liu, Xiangyu, et al.
Published: (2026)
by: Liu, Xiangyu, et al.
Published: (2026)
Video Editing for Audio-Visual Dubbing
by: Manela, Binyamin, et al.
Published: (2025)
by: Manela, Binyamin, et al.
Published: (2025)
LongCat-Video Technical Report
by: Meituan LongCat Team, et al.
Published: (2025)
by: Meituan LongCat Team, et al.
Published: (2025)
LongCat-Video-Avatar 1.5 Technical Report
by: Meituan LongCat Team, et al.
Published: (2026)
by: Meituan LongCat Team, et al.
Published: (2026)
JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion
by: Chen, Anthony, et al.
Published: (2026)
by: Chen, Anthony, et al.
Published: (2026)
Active Intelligence in Video Avatars via Closed-loop World Modeling
by: He, Xuanhua, et al.
Published: (2025)
by: He, Xuanhua, et al.
Published: (2025)
WildActor: Unconstrained Identity-Preserving Video Generation
by: Guo, Qin, et al.
Published: (2026)
by: Guo, Qin, et al.
Published: (2026)
LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models
by: Yu, Haojie, et al.
Published: (2025)
by: Yu, Haojie, et al.
Published: (2025)
A Comparative Study of Perceptual Quality Metrics for Audio-driven Talking Head Videos
by: Zhang, Weixia, et al.
Published: (2024)
by: Zhang, Weixia, et al.
Published: (2024)
MAViD: A Multimodal Framework for Audio-Visual Dialogue Understanding and Generation
by: Pang, Youxin, et al.
Published: (2025)
by: Pang, Youxin, et al.
Published: (2025)
MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
by: Zhong, Zhizhou, et al.
Published: (2025)
by: Zhong, Zhizhou, et al.
Published: (2025)
MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice
by: Yi, Hongwei, et al.
Published: (2025)
by: Yi, Hongwei, et al.
Published: (2025)
Infinite Video Understanding
by: Zhang, Dell, et al.
Published: (2025)
by: Zhang, Dell, et al.
Published: (2025)
MCDubber: Multimodal Context-Aware Expressive Video Dubbing
by: Zhao, Yuan, et al.
Published: (2024)
by: Zhao, Yuan, et al.
Published: (2024)
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
by: Tian, Wenjie, et al.
Published: (2025)
by: Tian, Wenjie, et al.
Published: (2025)
VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models
by: Sung-Bin, Kim, et al.
Published: (2025)
by: Sung-Bin, Kim, et al.
Published: (2025)
U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation
by: Deng, Xiang, et al.
Published: (2026)
by: Deng, Xiang, et al.
Published: (2026)
Length Aware Speech Translation for Video Dubbing
by: Chadha, Harveen Singh, et al.
Published: (2025)
by: Chadha, Harveen Singh, et al.
Published: (2025)
Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization
by: Cui, Chaoqun, et al.
Published: (2025)
by: Cui, Chaoqun, et al.
Published: (2025)
Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos
by: Feng, X., et al.
Published: (2026)
by: Feng, X., et al.
Published: (2026)
Audio-driven High-resolution Seamless Talking Head Video Editing via StyleGAN
by: Su, Jiacheng, et al.
Published: (2024)
by: Su, Jiacheng, et al.
Published: (2024)
StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation
by: Min, Dongchan, et al.
Published: (2022)
by: Min, Dongchan, et al.
Published: (2022)
SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model
by: Wang, Kaidi, et al.
Published: (2025)
by: Wang, Kaidi, et al.
Published: (2025)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
by: Li, Hebeizi, et al.
Published: (2026)
by: Li, Hebeizi, et al.
Published: (2026)
Identity-Preserving Video Dubbing Using Motion Warping
by: Liu, Runzhen, et al.
Published: (2025)
by: Liu, Runzhen, et al.
Published: (2025)
Text-driven Talking Face Synthesis by Reprogramming Audio-driven Models
by: Choi, Jeongsoo, et al.
Published: (2023)
by: Choi, Jeongsoo, et al.
Published: (2023)
GMTalker: Gaussian Mixture-based Audio-Driven Emotional Talking Video Portraits
by: Xia, Yibo, et al.
Published: (2023)
by: Xia, Yibo, et al.
Published: (2023)
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
by: Sahipjohn, Neha, et al.
Published: (2024)
by: Sahipjohn, Neha, et al.
Published: (2024)
StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation
by: Tu, Shuyuan, et al.
Published: (2025)
by: Tu, Shuyuan, et al.
Published: (2025)
In-Context Audio Control of Video Diffusion Transformers
by: Liu, Wenze, et al.
Published: (2025)
by: Liu, Wenze, et al.
Published: (2025)
CV-VAE: A Compatible Video VAE for Latent Generative Video Models
by: Zhao, Sijie, et al.
Published: (2024)
by: Zhao, Sijie, et al.
Published: (2024)
ANIM-400K: A Large-Scale Dataset for Automated End-To-End Dubbing of Video
by: Cai, Kevin, et al.
Published: (2024)
by: Cai, Kevin, et al.
Published: (2024)
PersonaTalk: Bring Attention to Your Persona in Visual Dubbing
by: Zhang, Longhao, et al.
Published: (2024)
by: Zhang, Longhao, et al.
Published: (2024)
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction
by: Zhao, Yuan, et al.
Published: (2024)
by: Zhao, Yuan, et al.
Published: (2024)
Similar Items
-
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
by: Kong, Zhe, et al.
Published: (2025) -
DAM-VSR: Disentanglement of Appearance and Motion for Video Super-Resolution
by: Kong, Zhe, et al.
Published: (2025) -
Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory
by: Wu, Ruiqi, et al.
Published: (2026) -
ZeroSmooth: Training-free Diffuser Adaptation for High Frame Rate Video Generation
by: Yang, Shaoshu, et al.
Published: (2024) -
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
by: Liu, Xiangyu, et al.
Published: (2026)