Multi-Modal Emotion Recognition by Text, Speech and Video Using Pretrained Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Shayaninasab, Minoo, Babaali, Bagher |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Persian Speech Emotion Recognition by Fine-Tuning Transformers
di: Shayaninasab, Minoo, et al.
Pubblicazione: (2024)
di: Shayaninasab, Minoo, et al.
Pubblicazione: (2024)
Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition
di: Muaz, Muhammad, et al.
Pubblicazione: (2024)
di: Muaz, Muhammad, et al.
Pubblicazione: (2024)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
di: Wu, Zhichao, et al.
Pubblicazione: (2025)
di: Wu, Zhichao, et al.
Pubblicazione: (2025)
Emotion Recognition and Generation: A Comprehensive Review of Face, Speech, and Text Modalities
di: Mobbs, Rebecca, et al.
Pubblicazione: (2025)
di: Mobbs, Rebecca, et al.
Pubblicazione: (2025)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
MGHFT: Multi-Granularity Hierarchical Fusion Transformer for Cross-Modal Sticker Emotion Recognition
di: Chen, Jian, et al.
Pubblicazione: (2025)
di: Chen, Jian, et al.
Pubblicazione: (2025)
Hybrid CNN-Transformer Architecture for Arabic Speech Emotion Recognition
di: Gheffari, Youcef Soufiane, et al.
Pubblicazione: (2026)
di: Gheffari, Youcef Soufiane, et al.
Pubblicazione: (2026)
Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition
di: Jain, Yash, et al.
Pubblicazione: (2024)
di: Jain, Yash, et al.
Pubblicazione: (2024)
Efficient Finetuning for Dimensional Speech Emotion Recognition in the Age of Transformers
di: Sampath, Aneesha, et al.
Pubblicazione: (2025)
di: Sampath, Aneesha, et al.
Pubblicazione: (2025)
CARAT: Contrastive Feature Reconstruction and Aggregation for Multi-Modal Multi-Label Emotion Recognition
di: Peng, Cheng, et al.
Pubblicazione: (2023)
di: Peng, Cheng, et al.
Pubblicazione: (2023)
MM-HSD: Multi-Modal Hate Speech Detection in Videos
di: Céspedes-Sarrias, Berta, et al.
Pubblicazione: (2025)
di: Céspedes-Sarrias, Berta, et al.
Pubblicazione: (2025)
Emotion Recognition Using Transformers with Masked Learning
di: Min, Seongjae, et al.
Pubblicazione: (2024)
di: Min, Seongjae, et al.
Pubblicazione: (2024)
Multi-Microphone Speech Emotion Recognition using the Hierarchical Token-semantic Audio Transformer Architecture
di: Cohen, Ohad, et al.
Pubblicazione: (2024)
di: Cohen, Ohad, et al.
Pubblicazione: (2024)
Plug-and-Play Emotion Graphs for Compositional Prompting in Zero-Shot Speech Emotion Recognition
di: Shi, Jiacheng, et al.
Pubblicazione: (2025)
di: Shi, Jiacheng, et al.
Pubblicazione: (2025)
Color-based Emotion Representation for Speech Emotion Recognition
di: Nagase, Ryotaro, et al.
Pubblicazione: (2026)
di: Nagase, Ryotaro, et al.
Pubblicazione: (2026)
A Pure Transformer Pretraining Framework on Text-attributed Graphs
di: Song, Yu, et al.
Pubblicazione: (2024)
di: Song, Yu, et al.
Pubblicazione: (2024)
Multi-Modal Character Localization and Extraction for Chinese Text Recognition
di: Li, Qilong, et al.
Pubblicazione: (2026)
di: Li, Qilong, et al.
Pubblicazione: (2026)
Speech Emotion Recognition via Entropy-Aware Score Selection
di: Chua, ChenYi, et al.
Pubblicazione: (2025)
di: Chua, ChenYi, et al.
Pubblicazione: (2025)
MATER: Multi-level Acoustic and Textual Emotion Representation for Interpretable Speech Emotion Recognition
di: Jon, Hyo Jin, et al.
Pubblicazione: (2025)
di: Jon, Hyo Jin, et al.
Pubblicazione: (2025)
Emotion Detection in Speech Using Lightweight and Transformer-Based Models: A Comparative and Ablation Study
di: Onyekwelu-Udoka, Lucky, et al.
Pubblicazione: (2025)
di: Onyekwelu-Udoka, Lucky, et al.
Pubblicazione: (2025)
Semantic Differentiation in Speech Emotion Recognition: Insights from Descriptive and Expressive Speech Roles
di: Guo, Rongchen, et al.
Pubblicazione: (2025)
di: Guo, Rongchen, et al.
Pubblicazione: (2025)
AIMDiT: Modality Augmentation and Interaction via Multimodal Dimension Transformation for Emotion Recognition in Conversations
di: Wu, Sheng, et al.
Pubblicazione: (2024)
di: Wu, Sheng, et al.
Pubblicazione: (2024)
Detecting Emotion Drift in Mental Health Text Using Pre-Trained Transformers
di: Sankpal, Shibani
Pubblicazione: (2025)
di: Sankpal, Shibani
Pubblicazione: (2025)
Speech Emotion Recognition Using MFCC Features and LSTM-Based Deep Learning Model
di: Oluwademilade, Adelekun, et al.
Pubblicazione: (2026)
di: Oluwademilade, Adelekun, et al.
Pubblicazione: (2026)
UDDETTS: Unifying Discrete and Dimensional Emotions for Controllable Emotional Text-to-Speech
di: Liu, Jiaxuan, et al.
Pubblicazione: (2025)
di: Liu, Jiaxuan, et al.
Pubblicazione: (2025)
Fuzzy Approach for Audio-Video Emotion Recognition in Computer Games for Children
di: Kozlov, Pavel, et al.
Pubblicazione: (2023)
di: Kozlov, Pavel, et al.
Pubblicazione: (2023)
Speech Emotion Recognition Using CNN and Its Use Case in Digital Healthcare
di: Nigar, Nishargo
Pubblicazione: (2024)
di: Nigar, Nishargo
Pubblicazione: (2024)
Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations
di: Chen, Jinming, et al.
Pubblicazione: (2025)
di: Chen, Jinming, et al.
Pubblicazione: (2025)
Centering Emotion Hotspots: Multimodal Local-Global Fusion and Cross-Modal Alignment for Emotion Recognition in Conversations
di: Liu, Yu, et al.
Pubblicazione: (2025)
di: Liu, Yu, et al.
Pubblicazione: (2025)
SER Evals: In-domain and Out-of-domain Benchmarking for Speech Emotion Recognition
di: Osman, Mohamed, et al.
Pubblicazione: (2024)
di: Osman, Mohamed, et al.
Pubblicazione: (2024)
Multi-scale Transformer-based Network for Emotion Recognition from Multi Physiological Signals
di: Vu, Tu, et al.
Pubblicazione: (2023)
di: Vu, Tu, et al.
Pubblicazione: (2023)
Multimodal Emotion Recognition with Vision-language Prompting and Modality Dropout
di: QI, Anbin, et al.
Pubblicazione: (2024)
di: QI, Anbin, et al.
Pubblicazione: (2024)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement
di: Yang, Jianing, et al.
Pubblicazione: (2025)
di: Yang, Jianing, et al.
Pubblicazione: (2025)
MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention
di: Jiao, Xinxin, et al.
Pubblicazione: (2024)
di: Jiao, Xinxin, et al.
Pubblicazione: (2024)
Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention
di: Wang, Cong, et al.
Pubblicazione: (2025)
di: Wang, Cong, et al.
Pubblicazione: (2025)
Multi-Modal Retrieval For Large Language Model Based Speech Recognition
di: Kolehmainen, Jari, et al.
Pubblicazione: (2024)
di: Kolehmainen, Jari, et al.
Pubblicazione: (2024)
EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition
di: Pan, Yu, et al.
Pubblicazione: (2023)
di: Pan, Yu, et al.
Pubblicazione: (2023)
Speech Emotion Recognition with Distilled Prosodic and Linguistic Affect Representations
di: Shome, Debaditya, et al.
Pubblicazione: (2023)
di: Shome, Debaditya, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Persian Speech Emotion Recognition by Fine-Tuning Transformers
di: Shayaninasab, Minoo, et al.
Pubblicazione: (2024) -
Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition
di: Muaz, Muhammad, et al.
Pubblicazione: (2024) -
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
di: Wu, Zhichao, et al.
Pubblicazione: (2025) -
Emotion Recognition and Generation: A Comprehensive Review of Face, Speech, and Text Modalities
di: Mobbs, Rebecca, et al.
Pubblicazione: (2025) -
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
di: Ma, Ziyang, et al.
Pubblicazione: (2023)