Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, Nanhan, Liu, Zhilei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
von: Wu, Kangyi, et al.
Veröffentlicht: (2025)
von: Wu, Kangyi, et al.
Veröffentlicht: (2025)
Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion Recognition
von: Xia, Haiying, et al.
Veröffentlicht: (2025)
von: Xia, Haiying, et al.
Veröffentlicht: (2025)
Memo2496: Expert-Annotated Dataset and Dual-View Adaptive Framework for Music Emotion Recognition
von: Li, Qilin, et al.
Veröffentlicht: (2025)
von: Li, Qilin, et al.
Veröffentlicht: (2025)
AMB-DSGDN: Adaptive Modality-Balanced Dynamic Semantic Graph Differential Network for Multimodal Emotion Recognition
von: Wang, Yunsheng, et al.
Veröffentlicht: (2026)
von: Wang, Yunsheng, et al.
Veröffentlicht: (2026)
Emotion-Aware Speech Generation with Character-Specific Voices for Comics
von: Qian, Zhiwen, et al.
Veröffentlicht: (2025)
von: Qian, Zhiwen, et al.
Veröffentlicht: (2025)
From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection
von: Liu, Ke, et al.
Veröffentlicht: (2026)
von: Liu, Ke, et al.
Veröffentlicht: (2026)
PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis
von: Xie, Yifan, et al.
Veröffentlicht: (2024)
von: Xie, Yifan, et al.
Veröffentlicht: (2024)
NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head Synthesis
von: Liu, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxing, et al.
Veröffentlicht: (2025)
Early Joint Learning of Emotion Information Makes MultiModal Model Understand You Better
von: Ge, Mengying, et al.
Veröffentlicht: (2024)
von: Ge, Mengying, et al.
Veröffentlicht: (2024)
DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
von: Hong, Fa-Ting, et al.
Veröffentlicht: (2024)
von: Hong, Fa-Ting, et al.
Veröffentlicht: (2024)
MusER: Musical Element-Based Regularization for Generating Symbolic Music with Emotion
von: Ji, Shulei, et al.
Veröffentlicht: (2023)
von: Ji, Shulei, et al.
Veröffentlicht: (2023)
MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition
von: Pan, Yu, et al.
Veröffentlicht: (2023)
von: Pan, Yu, et al.
Veröffentlicht: (2023)
D3PIA: A Discrete Denoising Diffusion Model for Piano Accompaniment Generation From Lead sheet
von: Choi, Eunjin, et al.
Veröffentlicht: (2026)
von: Choi, Eunjin, et al.
Veröffentlicht: (2026)
Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion
von: Deng, Zeyu, et al.
Veröffentlicht: (2025)
von: Deng, Zeyu, et al.
Veröffentlicht: (2025)
MotionBeat: Motion-Aligned Music Representation via Embodied Contrastive Learning and Bar-Equivariant Contact-Aware Encoding
von: Wang, Xuanchen, et al.
Veröffentlicht: (2025)
von: Wang, Xuanchen, et al.
Veröffentlicht: (2025)
EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
von: Li, Qifei, et al.
Veröffentlicht: (2024)
von: Li, Qifei, et al.
Veröffentlicht: (2024)
Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling
von: Fan, Congyi, et al.
Veröffentlicht: (2026)
von: Fan, Congyi, et al.
Veröffentlicht: (2026)
EARTalking: End-to-end GPT-style Autoregressive Talking Head Synthesis with Frame-wise Control
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
Bimodal Connection Attention Fusion for Speech Emotion Recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation
von: Sun, Jiahui, et al.
Veröffentlicht: (2025)
von: Sun, Jiahui, et al.
Veröffentlicht: (2025)
MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
von: Li, Haitian, et al.
Veröffentlicht: (2026)
von: Li, Haitian, et al.
Veröffentlicht: (2026)
EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
von: Ma, Ziyang, et al.
Veröffentlicht: (2024)
von: Ma, Ziyang, et al.
Veröffentlicht: (2024)
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
von: Sun, Yujia, et al.
Veröffentlicht: (2024)
von: Sun, Yujia, et al.
Veröffentlicht: (2024)
DRKF: Decoupled Representations with Knowledge Fusion for Multimodal Emotion Recognition
von: Jiang, Peiyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Peiyuan, et al.
Veröffentlicht: (2025)
LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the Wild
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
Stage-Adaptive Reliability Modeling for Continuous Valence-Arousal Estimation
von: Lee, Yubeen, et al.
Veröffentlicht: (2026)
von: Lee, Yubeen, et al.
Veröffentlicht: (2026)
EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction
von: Jing, Chong, et al.
Veröffentlicht: (2026)
von: Jing, Chong, et al.
Veröffentlicht: (2026)
S-PRESSO: Ultra Low Bitrate Sound Effect Compression With Diffusion Autoencoders And Offline Quantization
von: Lahrichi, Zineb, et al.
Veröffentlicht: (2026)
von: Lahrichi, Zineb, et al.
Veröffentlicht: (2026)
UniSRCodec: Unified and Low-Bitrate Single Codebook Codec with Sub-Band Reconstruction
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2026)
Variable-Length Audio Fingerprinting
von: Chen, Hongjie, et al.
Veröffentlicht: (2026)
von: Chen, Hongjie, et al.
Veröffentlicht: (2026)
Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training
von: You, Hong-Jie, et al.
Veröffentlicht: (2025)
von: You, Hong-Jie, et al.
Veröffentlicht: (2025)
Music Arena: Live Evaluation for Text-to-Music
von: Kim, Yonghyun, et al.
Veröffentlicht: (2025)
von: Kim, Yonghyun, et al.
Veröffentlicht: (2025)
MusicSwarm: Biologically Inspired Intelligence for Music Composition
von: Buehler, Markus J.
Veröffentlicht: (2025)
von: Buehler, Markus J.
Veröffentlicht: (2025)
Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic Music
von: Su, Hongju, et al.
Veröffentlicht: (2025)
von: Su, Hongju, et al.
Veröffentlicht: (2025)
SynthGuard: An Open Platform for Detecting AI-Generated Multimedia with Multimodal LLMs
von: Desai, Shail, et al.
Veröffentlicht: (2025)
von: Desai, Shail, et al.
Veröffentlicht: (2025)
Iterative Residual Cross-Attention Mechanism: An Integrated Approach for Audio-Visual Navigation Tasks
von: Zhang, Hailong, et al.
Veröffentlicht: (2025)
von: Zhang, Hailong, et al.
Veröffentlicht: (2025)
AUREXA-SE: Audio-Visual Unified Representation Exchange Architecture with Cross-Attention and Squeezeformer for Speech Enhancement
von: Sajid, M., et al.
Veröffentlicht: (2025)
von: Sajid, M., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
von: Wu, Kangyi, et al.
Veröffentlicht: (2025) -
Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion Recognition
von: Xia, Haiying, et al.
Veröffentlicht: (2025) -
Memo2496: Expert-Annotated Dataset and Dual-View Adaptive Framework for Music Emotion Recognition
von: Li, Qilin, et al.
Veröffentlicht: (2025) -
AMB-DSGDN: Adaptive Modality-Balanced Dynamic Semantic Graph Differential Network for Multimodal Emotion Recognition
von: Wang, Yunsheng, et al.
Veröffentlicht: (2026) -
Emotion-Aware Speech Generation with Character-Specific Voices for Comics
von: Qian, Zhiwen, et al.
Veröffentlicht: (2025)