TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Xiangyu, Gao, Feng, Zhang, Xiaomei, Zhang, Yong, Wei, Xiaoming, Lei, Zhen, Zhu, Xiangyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
von: Ye, Zhen, et al.
Veröffentlicht: (2026)
von: Ye, Zhen, et al.
Veröffentlicht: (2026)
Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence
von: Liao, Junchao, et al.
Veröffentlicht: (2026)
von: Liao, Junchao, et al.
Veröffentlicht: (2026)
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
von: Li, Chunyu, et al.
Veröffentlicht: (2026)
von: Li, Chunyu, et al.
Veröffentlicht: (2026)
From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection
von: Liu, Ke, et al.
Veröffentlicht: (2026)
von: Liu, Ke, et al.
Veröffentlicht: (2026)
PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis
von: Xie, Yifan, et al.
Veröffentlicht: (2024)
von: Xie, Yifan, et al.
Veröffentlicht: (2024)
CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
von: Wu, Kangyi, et al.
Veröffentlicht: (2025)
von: Wu, Kangyi, et al.
Veröffentlicht: (2025)
Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior Distillation
von: Shen, Nanhan, et al.
Veröffentlicht: (2026)
von: Shen, Nanhan, et al.
Veröffentlicht: (2026)
Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation
von: Kang, Fang, et al.
Veröffentlicht: (2025)
von: Kang, Fang, et al.
Veröffentlicht: (2025)
Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the Wild
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
GAIA: Zero-shot Talking Avatar Generation
von: He, Tianyu, et al.
Veröffentlicht: (2023)
von: He, Tianyu, et al.
Veröffentlicht: (2023)
SonoWorld: From One Image to a 3D Audio-Visual Scene
von: Jin, Derong, et al.
Veröffentlicht: (2026)
von: Jin, Derong, et al.
Veröffentlicht: (2026)
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
von: Shi, Weiyan, et al.
Veröffentlicht: (2025)
von: Shi, Weiyan, et al.
Veröffentlicht: (2025)
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
AudioStory: Generating Long-Form Narrative Audio with Large Language Models
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
EARTalking: End-to-end GPT-style Autoregressive Talking Head Synthesis with Frame-wise Control
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head Synthesis
von: Liu, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxing, et al.
Veröffentlicht: (2025)
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
von: Liu, Kai, et al.
Veröffentlicht: (2026)
von: Liu, Kai, et al.
Veröffentlicht: (2026)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
XGC-AVis: Towards Audio-Visual Content Understanding with a Multi-Agent Collaborative System
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
von: Guo, Xinyue, et al.
Veröffentlicht: (2025)
von: Guo, Xinyue, et al.
Veröffentlicht: (2025)
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
von: Tian, Zeyue, et al.
Veröffentlicht: (2026)
von: Tian, Zeyue, et al.
Veröffentlicht: (2026)
3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation
von: Li, Yaoru, et al.
Veröffentlicht: (2025)
von: Li, Yaoru, et al.
Veröffentlicht: (2025)
HDA-SELD: Hierarchical Cross-Modal Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
von: Yang, Jianxuan, et al.
Veröffentlicht: (2026)
von: Yang, Jianxuan, et al.
Veröffentlicht: (2026)
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
von: Wang, Le, et al.
Veröffentlicht: (2025)
von: Wang, Le, et al.
Veröffentlicht: (2025)
FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation
von: Tan, Weiting, et al.
Veröffentlicht: (2026)
von: Tan, Weiting, et al.
Veröffentlicht: (2026)
Generative Audio Extension and Morphing
von: Seetharaman, Prem, et al.
Veröffentlicht: (2026)
von: Seetharaman, Prem, et al.
Veröffentlicht: (2026)
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
von: Liu, Tengfei, et al.
Veröffentlicht: (2026)
von: Liu, Tengfei, et al.
Veröffentlicht: (2026)
Video-to-Audio Generation with Hidden Alignment
von: Xu, Manjie, et al.
Veröffentlicht: (2024)
von: Xu, Manjie, et al.
Veröffentlicht: (2024)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
von: Ling, Jun, et al.
Veröffentlicht: (2024)
von: Ling, Jun, et al.
Veröffentlicht: (2024)
Weakly-supervised Audio Temporal Forgery Localization via Progressive Audio-language Co-learning Network
von: Wu, Junyan, et al.
Veröffentlicht: (2025)
von: Wu, Junyan, et al.
Veröffentlicht: (2025)
OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
von: Su, Yaofeng, et al.
Veröffentlicht: (2026)
von: Su, Yaofeng, et al.
Veröffentlicht: (2026)
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text
von: Pian, Weiguo, et al.
Veröffentlicht: (2026)
von: Pian, Weiguo, et al.
Veröffentlicht: (2026)
DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
von: Hong, Fa-Ting, et al.
Veröffentlicht: (2024)
von: Hong, Fa-Ting, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
von: Li, Hebeizi, et al.
Veröffentlicht: (2026) -
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025) -
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
von: Ye, Zhen, et al.
Veröffentlicht: (2026) -
Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence
von: Liao, Junchao, et al.
Veröffentlicht: (2026) -
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
von: Li, Chunyu, et al.
Veröffentlicht: (2026)