A Multimodal Emotion Recognition System: Integrating Facial Expressions, Body Movement, Speech, and Spoken Language
Fuente:
arXiv
Salvato in:
| Autore principale: | Kraack, Kris |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition
di: Sun, Licai, et al.
Pubblicazione: (2024)
di: Sun, Licai, et al.
Pubblicazione: (2024)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
di: Nishida, Naoto, et al.
Pubblicazione: (2025)
di: Nishida, Naoto, et al.
Pubblicazione: (2025)
Towards Reliable Large Audio Language Model
di: Ma, Ziyang, et al.
Pubblicazione: (2025)
di: Ma, Ziyang, et al.
Pubblicazione: (2025)
MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes
di: Chen, Maximillian, et al.
Pubblicazione: (2026)
di: Chen, Maximillian, et al.
Pubblicazione: (2026)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
di: Cai, Zhuojiang, et al.
Pubblicazione: (2024)
di: Cai, Zhuojiang, et al.
Pubblicazione: (2024)
Soundify: Matching Sound Effects to Video
di: Lin, David Chuan-En, et al.
Pubblicazione: (2021)
di: Lin, David Chuan-En, et al.
Pubblicazione: (2021)
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
di: Tan, Weiting, et al.
Pubblicazione: (2025)
di: Tan, Weiting, et al.
Pubblicazione: (2025)
Capturing Cancer as Music: Cancer Mechanisms Expressed through Musification
di: Hnatyshyn, Rostyslav, et al.
Pubblicazione: (2024)
di: Hnatyshyn, Rostyslav, et al.
Pubblicazione: (2024)
Assessing the Viability of Wave Field Synthesis in VR-Based Cognitive Research
di: Kahl, Benjamin
Pubblicazione: (2025)
di: Kahl, Benjamin
Pubblicazione: (2025)
Creating Aesthetic Sonifications on the Web with SIREN
di: Peng, Tristan, et al.
Pubblicazione: (2024)
di: Peng, Tristan, et al.
Pubblicazione: (2024)
MR-DAW: Towards Collaborative Digital Audio Workstations in Mixed Reality
di: Hopkins, Torin, et al.
Pubblicazione: (2026)
di: Hopkins, Torin, et al.
Pubblicazione: (2026)
VidTune: Creating Video Soundtracks with Generative Music and Contextual Thumbnails
di: Huh, Mina, et al.
Pubblicazione: (2026)
di: Huh, Mina, et al.
Pubblicazione: (2026)
NeoLightning: A Modern Reimagination of Gesture-Based Sound Design
di: Kim, Yonghyun, et al.
Pubblicazione: (2025)
di: Kim, Yonghyun, et al.
Pubblicazione: (2025)
MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction
di: He, Jiajun, et al.
Pubblicazione: (2024)
di: He, Jiajun, et al.
Pubblicazione: (2024)
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
di: Fu, Yu-Kuan, et al.
Pubblicazione: (2024)
di: Fu, Yu-Kuan, et al.
Pubblicazione: (2024)
Are Expressions for Music Emotions the Same Across Cultures?
di: Celen, Elif, et al.
Pubblicazione: (2025)
di: Celen, Elif, et al.
Pubblicazione: (2025)
MLLM-based Speech Recognition: When and How is Multimodality Beneficial?
di: Guan, Yiwen, et al.
Pubblicazione: (2025)
di: Guan, Yiwen, et al.
Pubblicazione: (2025)
A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation
di: Min, Anna, et al.
Pubblicazione: (2025)
di: Min, Anna, et al.
Pubblicazione: (2025)
SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition
di: Cheng, Zebang, et al.
Pubblicazione: (2024)
di: Cheng, Zebang, et al.
Pubblicazione: (2024)
Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques
di: Li, Yuanchao, et al.
Pubblicazione: (2024)
di: Li, Yuanchao, et al.
Pubblicazione: (2024)
Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs
di: Cappellazzo, Umberto, et al.
Pubblicazione: (2025)
di: Cappellazzo, Umberto, et al.
Pubblicazione: (2025)
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
di: Yeo, Jeong Hun, et al.
Pubblicazione: (2025)
di: Yeo, Jeong Hun, et al.
Pubblicazione: (2025)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
di: Cheng, Xize, et al.
Pubblicazione: (2025)
di: Cheng, Xize, et al.
Pubblicazione: (2025)
G-STAR: End-to-End Global Speaker-Tracking Attributed Recognition
di: Peng, Jing, et al.
Pubblicazione: (2026)
di: Peng, Jing, et al.
Pubblicazione: (2026)
Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models
di: Dietrich, Juergen
Pubblicazione: (2026)
di: Dietrich, Juergen
Pubblicazione: (2026)
Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness
di: Yang, Sicheng, et al.
Pubblicazione: (2024)
di: Yang, Sicheng, et al.
Pubblicazione: (2024)
The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
di: Zhao, Zhixian, et al.
Pubblicazione: (2026)
di: Zhao, Zhixian, et al.
Pubblicazione: (2026)
Flowers Revisited: A Preliminary Replication of Flowers et al. 1997
di: Enge, Kajetan, et al.
Pubblicazione: (2024)
di: Enge, Kajetan, et al.
Pubblicazione: (2024)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
di: Cappellazzo, Umberto, et al.
Pubblicazione: (2024)
di: Cappellazzo, Umberto, et al.
Pubblicazione: (2024)
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
di: Goncalves, Lucas, et al.
Pubblicazione: (2024)
di: Goncalves, Lucas, et al.
Pubblicazione: (2024)
Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
di: Yeo, Jeong Hun, et al.
Pubblicazione: (2025)
di: Yeo, Jeong Hun, et al.
Pubblicazione: (2025)
Addressing Emotion Bias in Music Emotion Recognition and Generation with Frechet Audio Distance
di: Li, Yuanchao, et al.
Pubblicazione: (2024)
di: Li, Yuanchao, et al.
Pubblicazione: (2024)
A Joint Cross-Attention Model for Audio-Visual Fusion in Dimensional Emotion Recognition
di: Praveen, R. Gnana, et al.
Pubblicazione: (2022)
di: Praveen, R. Gnana, et al.
Pubblicazione: (2022)
Workflow-Based Evaluation of Music Generation Systems
di: Dadman, Shayan, et al.
Pubblicazione: (2025)
di: Dadman, Shayan, et al.
Pubblicazione: (2025)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
di: Chen, Youjun, et al.
Pubblicazione: (2025)
di: Chen, Youjun, et al.
Pubblicazione: (2025)
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation
di: Izzati, Fathinah, et al.
Pubblicazione: (2025)
di: Izzati, Fathinah, et al.
Pubblicazione: (2025)
Design and Development of Laughter Recognition System Based on Multimodal Fusion and Deep Learning
di: Zhao, Fuzheng, et al.
Pubblicazione: (2024)
di: Zhao, Fuzheng, et al.
Pubblicazione: (2024)
Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
di: Zhao, Zhixian, et al.
Pubblicazione: (2024)
di: Zhao, Zhixian, et al.
Pubblicazione: (2024)
Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training
di: He, Jianfeng, et al.
Pubblicazione: (2023)
di: He, Jianfeng, et al.
Pubblicazione: (2023)
Documenti analoghi
-
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
di: Zhou, Dongliang, et al.
Pubblicazione: (2025) -
HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition
di: Sun, Licai, et al.
Pubblicazione: (2024) -
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
di: Nishida, Naoto, et al.
Pubblicazione: (2025) -
Towards Reliable Large Audio Language Model
di: Ma, Ziyang, et al.
Pubblicazione: (2025) -
MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes
di: Chen, Maximillian, et al.
Pubblicazione: (2026)