EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian Deformation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Tianheng, Yu, Yinfeng, Wang, Liejun, Sun, Fuchun, Zheng, Wendong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning
von: Zhu, Tao, et al.
Veröffentlicht: (2025)
von: Zhu, Tao, et al.
Veröffentlicht: (2025)
PGSTalker: Real-Time Audio-Driven Talking Head Generation via 3D Gaussian Splatting with Pixel-Aware Density Control
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025)
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025)
ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations
von: Wang, Kexue, et al.
Veröffentlicht: (2026)
von: Wang, Kexue, et al.
Veröffentlicht: (2026)
Leveraging Label Potential for Enhanced Multimodal Emotion Recognition
von: Shao, Xuechun, et al.
Veröffentlicht: (2025)
von: Shao, Xuechun, et al.
Veröffentlicht: (2025)
MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
Learning Frame-Wise Emotion Intensity for Audio-Driven Talking-Head Generation
von: Xu, Jingyi, et al.
Veröffentlicht: (2024)
von: Xu, Jingyi, et al.
Veröffentlicht: (2024)
Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
Magnitude-Phase Dual-Path Speech Enhancement Network based on Self-Supervised Embedding and Perceptual Contrast Stretch Boosting
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2025)
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2025)
AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis
von: Cao, Yubing, et al.
Veröffentlicht: (2025)
von: Cao, Yubing, et al.
Veröffentlicht: (2025)
Audio Spatially-Guided Fusion for Audio-Visual Navigation
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis
von: Xie, Yifan, et al.
Veröffentlicht: (2024)
von: Xie, Yifan, et al.
Veröffentlicht: (2024)
Spatial-Aware Conditioned Fusion for Audio-Visual Navigation
von: Wu, Shaohang, et al.
Veröffentlicht: (2026)
von: Wu, Shaohang, et al.
Veröffentlicht: (2026)
PCQ: Emotion Recognition in Speech via Progressive Channel Querying
von: Wang, Xincheng, et al.
Veröffentlicht: (2024)
von: Wang, Xincheng, et al.
Veröffentlicht: (2024)
Reliability-Aware Geometric Fusion for Robust Audio-Visual Navigation
von: Liu, Teng, et al.
Veröffentlicht: (2026)
von: Liu, Teng, et al.
Veröffentlicht: (2026)
Generalizable Audio-Visual Navigation via Binaural Difference Attention and Action Transition Prediction
von: Li, Jia, et al.
Veröffentlicht: (2026)
von: Li, Jia, et al.
Veröffentlicht: (2026)
LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the Wild
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
von: Hong, Fa-Ting, et al.
Veröffentlicht: (2024)
von: Hong, Fa-Ting, et al.
Veröffentlicht: (2024)
Audio-Driven Reinforcement Learning for Head-Orientation in Naturalistic Environments
von: Ledder, Wessel, et al.
Veröffentlicht: (2024)
von: Ledder, Wessel, et al.
Veröffentlicht: (2024)
Audio-Plane: Audio Factorization Plane Gaussian Splatting for Real-Time Talking Head Synthesis
von: Shen, Shuai, et al.
Veröffentlicht: (2025)
von: Shen, Shuai, et al.
Veröffentlicht: (2025)
Hyperdimensional Intelligent Sensing for Efficient Real-Time Audio Processing on Extreme Edge
von: Yun, Sanggeon, et al.
Veröffentlicht: (2025)
von: Yun, Sanggeon, et al.
Veröffentlicht: (2025)
AudioGAN: A Compact and Efficient Framework for Real-Time High-Fidelity Text-to-Audio Generation
von: Chung, HaeChun
Veröffentlicht: (2025)
von: Chung, HaeChun
Veröffentlicht: (2025)
Modality-Invariant Bidirectional Temporal Representation Distillation Network for Missing Multimodal Sentiment Analysis
von: Wang, Xincheng, et al.
Veröffentlicht: (2025)
von: Wang, Xincheng, et al.
Veröffentlicht: (2025)
Go witheFlow: Real-time Emotion Driven Audio Effects Modulation
von: Dervakos, Edmund, et al.
Veröffentlicht: (2025)
von: Dervakos, Edmund, et al.
Veröffentlicht: (2025)
READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation
von: Wang, Haotian, et al.
Veröffentlicht: (2025)
von: Wang, Haotian, et al.
Veröffentlicht: (2025)
Two-stage Audio-Visual Target Speaker Extraction System for Real-Time Processing On Edge Device
von: Li, Zixuan, et al.
Veröffentlicht: (2025)
von: Li, Zixuan, et al.
Veröffentlicht: (2025)
Warm Chat: Diffuse Emotion-aware Interactive Talking Head Avatar with Tree-Structured Guidance
von: Yang, Haijie, et al.
Veröffentlicht: (2025)
von: Yang, Haijie, et al.
Veröffentlicht: (2025)
Heterogeneous Space Fusion and Dual-Dimension Attention: A New Paradigm for Speech Enhancement
von: Zheng, Tao, et al.
Veröffentlicht: (2024)
von: Zheng, Tao, et al.
Veröffentlicht: (2024)
MiDashengLM: Efficient Audio Understanding with General Audio Captions
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
Audio Explanation Synthesis with Generative Foundation Models
von: Akman, Alican, et al.
Veröffentlicht: (2024)
von: Akman, Alican, et al.
Veröffentlicht: (2024)
FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head Models
von: Aneja, Shivangi, et al.
Veröffentlicht: (2023)
von: Aneja, Shivangi, et al.
Veröffentlicht: (2023)
FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio Generation
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face Animation
von: Zhou, Xukun, et al.
Veröffentlicht: (2024)
von: Zhou, Xukun, et al.
Veröffentlicht: (2024)
Efficient Autoregressive Audio Modeling via Next-Scale Prediction
von: Qiu, Kai, et al.
Veröffentlicht: (2024)
von: Qiu, Kai, et al.
Veröffentlicht: (2024)
Dual Audio-Centric Modality Coupling for Talking Head Generation
von: Fu, Ao, et al.
Veröffentlicht: (2025)
von: Fu, Ao, et al.
Veröffentlicht: (2025)
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
Generalized Source Tracing: Detecting Novel Audio Deepfake Algorithm with Real Emphasis and Fake Dispersion Strategy
von: Xie, Yuankun, et al.
Veröffentlicht: (2024)
von: Xie, Yuankun, et al.
Veröffentlicht: (2024)
SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake Detection
von: Zhu, Yi, et al.
Veröffentlicht: (2024)
von: Zhu, Yi, et al.
Veröffentlicht: (2024)
AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning
von: Zhu, Tao, et al.
Veröffentlicht: (2025) -
PGSTalker: Real-Time Audio-Driven Talking Head Generation via 3D Gaussian Splatting with Pixel-Aware Density Control
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025) -
ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations
von: Wang, Kexue, et al.
Veröffentlicht: (2026) -
Leveraging Label Potential for Enhanced Multimodal Emotion Recognition
von: Shao, Xuechun, et al.
Veröffentlicht: (2025) -
MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)