Emotional Face-to-Speech
Fuente:
arXiv
Guardado en:
| Autores principales: | Ye, Jiaxin, Cao, Boyuan, Shan, Hongming |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Shushing! Let's Imagine an Authentic Speech from the Silent Video
por: Ye, Jiaxin, et al.
Publicado: (2025)
por: Ye, Jiaxin, et al.
Publicado: (2025)
EmoTalker: Emotionally Editable Talking Face Generation via Diffusion Model
por: Zhang, Bingyuan, et al.
Publicado: (2024)
por: Zhang, Bingyuan, et al.
Publicado: (2024)
Emotional Vietnamese Speech-Based Depression Diagnosis Using Dynamic Attention Mechanism
por: D., Quang-Anh N., et al.
Publicado: (2024)
por: D., Quang-Anh N., et al.
Publicado: (2024)
UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech Generation
por: Wang, Jinting, et al.
Publicado: (2025)
por: Wang, Jinting, et al.
Publicado: (2025)
Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation
por: Kang, Fang, et al.
Publicado: (2025)
por: Kang, Fang, et al.
Publicado: (2025)
ESARM: 3D Emotional Speech-to-Animation via Reward Model from Automatically-Ranked Demonstrations
por: Zhang, Xulong, et al.
Publicado: (2024)
por: Zhang, Xulong, et al.
Publicado: (2024)
Hear Your Face: Face-based voice conversion with F0 estimation
por: Lee, Jaejun, et al.
Publicado: (2024)
por: Lee, Jaejun, et al.
Publicado: (2024)
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
por: Fu, Chaoyou, et al.
Publicado: (2025)
por: Fu, Chaoyou, et al.
Publicado: (2025)
A Low-rank Matching Attention based Cross-modal Feature Fusion Method for Conversational Emotion Recognition
por: Shou, Yuntao, et al.
Publicado: (2023)
por: Shou, Yuntao, et al.
Publicado: (2023)
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
por: Choi, Jeongsoo, et al.
Publicado: (2024)
por: Choi, Jeongsoo, et al.
Publicado: (2024)
Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech Representation
por: Kim, Minsu, et al.
Publicado: (2024)
por: Kim, Minsu, et al.
Publicado: (2024)
Face-voice Association in Multilingual Environments (FAME) Challenge 2024 Evaluation Plan
por: Saeed, Muhammad Saad, et al.
Publicado: (2024)
por: Saeed, Muhammad Saad, et al.
Publicado: (2024)
Audio-Assisted Face Video Restoration with Temporal and Identity Complementary Learning
por: Cao, Yuqin, et al.
Publicado: (2025)
por: Cao, Yuqin, et al.
Publicado: (2025)
Enhancing CTC-Based Visual Speech Recognition
por: Laux, Hendrik, et al.
Publicado: (2024)
por: Laux, Hendrik, et al.
Publicado: (2024)
Multimodal Emotion Recognition and Sentiment Analysis in Multi-Party Conversation Contexts
por: Farhadipour, Aref, et al.
Publicado: (2025)
por: Farhadipour, Aref, et al.
Publicado: (2025)
Input Conditioned Layer Dropping in Speech Foundation Models
por: Hannan, Abdul, et al.
Publicado: (2025)
por: Hannan, Abdul, et al.
Publicado: (2025)
Recursive Joint Cross-Modal Attention for Multimodal Fusion in Dimensional Emotion Recognition
por: Praveen, R. Gnana, et al.
Publicado: (2024)
por: Praveen, R. Gnana, et al.
Publicado: (2024)
Decoding Emotions: Unveiling Facial Expressions through Acoustic Sensing with Contrastive Attention
por: Wang, Guangjing, et al.
Publicado: (2024)
por: Wang, Guangjing, et al.
Publicado: (2024)
Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes
por: Ryu, Hyeonggon, et al.
Publicado: (2025)
por: Ryu, Hyeonggon, et al.
Publicado: (2025)
Spiking Structured State Space Model for Monaural Speech Enhancement
por: Du, Yu, et al.
Publicado: (2023)
por: Du, Yu, et al.
Publicado: (2023)
From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech
por: Kim, Ji-Hoon, et al.
Publicado: (2025)
por: Kim, Ji-Hoon, et al.
Publicado: (2025)
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
por: Cappellazzo, Umberto, et al.
Publicado: (2025)
por: Cappellazzo, Umberto, et al.
Publicado: (2025)
CNVSRC 2024: The Second Chinese Continuous Visual Speech Recognition Challenge
por: Liu, Zehua, et al.
Publicado: (2025)
por: Liu, Zehua, et al.
Publicado: (2025)
From Detection to Correction: Backdoor-Resilient Face Recognition via Vision-Language Trigger Detection and Noise-Based Neutralization
por: Wahida, Farah, et al.
Publicado: (2025)
por: Wahida, Farah, et al.
Publicado: (2025)
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition
por: Rouditchenko, Andrew, et al.
Publicado: (2025)
por: Rouditchenko, Andrew, et al.
Publicado: (2025)
Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
por: Cappellazzo, Umberto, et al.
Publicado: (2025)
por: Cappellazzo, Umberto, et al.
Publicado: (2025)
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
por: Anand, et al.
Publicado: (2025)
por: Anand, et al.
Publicado: (2025)
Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech Recognition
por: Liu, Lei, et al.
Publicado: (2024)
por: Liu, Lei, et al.
Publicado: (2024)
SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer
por: Park, Young-Hu, et al.
Publicado: (2025)
por: Park, Young-Hu, et al.
Publicado: (2025)
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling
por: Pham, Long-Khanh, et al.
Publicado: (2025)
por: Pham, Long-Khanh, et al.
Publicado: (2025)
RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation
por: Pegg, Samuel, et al.
Publicado: (2023)
por: Pegg, Samuel, et al.
Publicado: (2023)
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
por: Rouditchenko, Andrew, et al.
Publicado: (2024)
por: Rouditchenko, Andrew, et al.
Publicado: (2024)
VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language Models
por: Hu, Rui, et al.
Publicado: (2025)
por: Hu, Rui, et al.
Publicado: (2025)
CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
por: Chen, Xueyuan, et al.
Publicado: (2024)
por: Chen, Xueyuan, et al.
Publicado: (2024)
Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation
por: Li, Hao, et al.
Publicado: (2025)
por: Li, Hao, et al.
Publicado: (2025)
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing
por: Liu, Zehua, et al.
Publicado: (2025)
por: Liu, Zehua, et al.
Publicado: (2025)
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
por: Cappellazzo, Umberto, et al.
Publicado: (2026)
por: Cappellazzo, Umberto, et al.
Publicado: (2026)
United we stand, Divided we fall: Handling Weak Complementary Relationships for Audio-Visual Emotion Recognition in Valence-Arousal Space
por: Praveen, R. Gnana, et al.
Publicado: (2025)
por: Praveen, R. Gnana, et al.
Publicado: (2025)
NaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis with Differential Digital Signal Processing
por: Liang, Yifan, et al.
Publicado: (2025)
por: Liang, Yifan, et al.
Publicado: (2025)
DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation
por: Zhang, Haomin, et al.
Publicado: (2025)
por: Zhang, Haomin, et al.
Publicado: (2025)
Ejemplares similares
-
Shushing! Let's Imagine an Authentic Speech from the Silent Video
por: Ye, Jiaxin, et al.
Publicado: (2025) -
EmoTalker: Emotionally Editable Talking Face Generation via Diffusion Model
por: Zhang, Bingyuan, et al.
Publicado: (2024) -
Emotional Vietnamese Speech-Based Depression Diagnosis Using Dynamic Attention Mechanism
por: D., Quang-Anh N., et al.
Publicado: (2024) -
UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech Generation
por: Wang, Jinting, et al.
Publicado: (2025) -
Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation
por: Kang, Fang, et al.
Publicado: (2025)