Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zhicheng, Wang, Lei, Zhang, Yu, Gao, Yongsheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
von: Ling, Jun, et al.
Veröffentlicht: (2024)
von: Ling, Jun, et al.
Veröffentlicht: (2024)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
VineetVC: Adaptive Video Conferencing Under Severe Bandwidth Constraints Using Audio-Driven Talking-Head Reconstruction
von: Rakesh, Vineet Kumar, et al.
Veröffentlicht: (2026)
von: Rakesh, Vineet Kumar, et al.
Veröffentlicht: (2026)
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
von: Liu, Xiangyu, et al.
Veröffentlicht: (2026)
von: Liu, Xiangyu, et al.
Veröffentlicht: (2026)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework
von: Liu, Ke, et al.
Veröffentlicht: (2026)
von: Liu, Ke, et al.
Veröffentlicht: (2026)
Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation
von: Huang, Feizhen, et al.
Veröffentlicht: (2025)
von: Huang, Feizhen, et al.
Veröffentlicht: (2025)
EARTalking: End-to-end GPT-style Autoregressive Talking Head Synthesis with Frame-wise Control
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation
von: Qi, Xingqun, et al.
Veröffentlicht: (2023)
von: Qi, Xingqun, et al.
Veröffentlicht: (2023)
Talking Head Generation Driven by Speech-Related Facial Action Units and Audio- Based on Multimodal Representation Fusion
von: Chen, Sen, et al.
Veröffentlicht: (2022)
von: Chen, Sen, et al.
Veröffentlicht: (2022)
GaussianTalker: Real-Time High-Fidelity Talking Head Synthesis with Audio-Driven 3D Gaussian Splatting
von: Cho, Kyusun, et al.
Veröffentlicht: (2024)
von: Cho, Kyusun, et al.
Veröffentlicht: (2024)
Neuron Abandoning Attention Flow: Visual Explanation of Dynamics inside CNN Models
von: Liao, Yi, et al.
Veröffentlicht: (2024)
von: Liao, Yi, et al.
Veröffentlicht: (2024)
FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation
von: Tan, Weiting, et al.
Veröffentlicht: (2026)
von: Tan, Weiting, et al.
Veröffentlicht: (2026)
Audio Visual Segmentation Through Text Embeddings
von: Lee, Kyungbok, et al.
Veröffentlicht: (2025)
von: Lee, Kyungbok, et al.
Veröffentlicht: (2025)
Audio-Guided Visual Perception for Audio-Visual Navigation
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
von: Tian, Zeyue, et al.
Veröffentlicht: (2026)
von: Tian, Zeyue, et al.
Veröffentlicht: (2026)
Taming Modality Entanglement in Continual Audio-Visual Segmentation
von: Hong, Yuyang, et al.
Veröffentlicht: (2025)
von: Hong, Yuyang, et al.
Veröffentlicht: (2025)
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
von: Yao, Lei, et al.
Veröffentlicht: (2025)
von: Yao, Lei, et al.
Veröffentlicht: (2025)
Is It Really You? Exploring Biometric Verification Scenarios in Photorealistic Talking-Head Avatar Videos
von: Pedrouzo-Rodriguez, Laura, et al.
Veröffentlicht: (2025)
von: Pedrouzo-Rodriguez, Laura, et al.
Veröffentlicht: (2025)
Beyond Audio and Pose: A General-Purpose Framework for Video Synchronization
von: Shin, Yosub, et al.
Veröffentlicht: (2025)
von: Shin, Yosub, et al.
Veröffentlicht: (2025)
Controllable Audio-Visual Viewpoint Generation from 360° Spatial Information
von: Marinoni, Christian, et al.
Veröffentlicht: (2025)
von: Marinoni, Christian, et al.
Veröffentlicht: (2025)
LLM-EvRep: Learning an LLM-Compatible Event Representation Using a Self-Supervised Framework
von: Yu, Zongyou, et al.
Veröffentlicht: (2025)
von: Yu, Zongyou, et al.
Veröffentlicht: (2025)
Style-Preserving Lip Sync via Audio-Aware Style Reference
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
Apollo: Unified Multi-Task Audio-Video Joint Generation
von: Wang, Jun, et al.
Veröffentlicht: (2026)
von: Wang, Jun, et al.
Veröffentlicht: (2026)
Discover Your Neighbors: Advanced Stable Test-Time Adaptation in Dynamic World
von: Jiang, Qinting, et al.
Veröffentlicht: (2024)
von: Jiang, Qinting, et al.
Veröffentlicht: (2024)
EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers
von: Flynn, John, et al.
Veröffentlicht: (2026)
von: Flynn, John, et al.
Veröffentlicht: (2026)
Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation
von: Liang, Yingshan, et al.
Veröffentlicht: (2025)
von: Liang, Yingshan, et al.
Veröffentlicht: (2025)
AutoAWG: Adverse Weather Generation with Adaptive Multi-Controls for Automotive Videos
von: Hu, Jiagao, et al.
Veröffentlicht: (2026)
von: Hu, Jiagao, et al.
Veröffentlicht: (2026)
DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression
von: Li, Bingzhou, et al.
Veröffentlicht: (2026)
von: Li, Bingzhou, et al.
Veröffentlicht: (2026)
Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
von: Cai, Dongnuan, et al.
Veröffentlicht: (2026)
von: Cai, Dongnuan, et al.
Veröffentlicht: (2026)
Diffusion Models for Joint Audio-Video Generation
von: La Torre, Alejandro Paredes
Veröffentlicht: (2026)
von: La Torre, Alejandro Paredes
Veröffentlicht: (2026)
FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion Model
von: Yao, Ziyu, et al.
Veröffentlicht: (2024)
von: Yao, Ziyu, et al.
Veröffentlicht: (2024)
Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR
von: Zhang, Yulong, et al.
Veröffentlicht: (2025)
von: Zhang, Yulong, et al.
Veröffentlicht: (2025)
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
High-fidelity and Lip-synced Talking Face Synthesis via Landmark-based Diffusion Model
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
Decoupled Audio-Visual Dataset Distillation
von: Li, Wenyuan, et al.
Veröffentlicht: (2025)
von: Li, Wenyuan, et al.
Veröffentlicht: (2025)
TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
von: Huang, Victor Shea-Jay, et al.
Veröffentlicht: (2025)
von: Huang, Victor Shea-Jay, et al.
Veröffentlicht: (2025)
UniVid: Pyramid Diffusion Model for High Quality Video Generation
von: Xiao, Xinyu, et al.
Veröffentlicht: (2026)
von: Xiao, Xinyu, et al.
Veröffentlicht: (2026)
ScaleDreamer: Scalable Text-to-3D Synthesis with Asynchronous Score Distillation
von: Ma, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Ma, Zhiyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
von: Ling, Jun, et al.
Veröffentlicht: (2024) -
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025) -
VineetVC: Adaptive Video Conferencing Under Severe Bandwidth Constraints Using Audio-Driven Talking-Head Reconstruction
von: Rakesh, Vineet Kumar, et al.
Veröffentlicht: (2026) -
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
von: Liu, Xiangyu, et al.
Veröffentlicht: (2026) -
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
von: Gan, Qijun, et al.
Veröffentlicht: (2025)