Speech Emotion Recognition Via CNN-Transformer and Multidimensional Attention Mechanism
Fuente:
arXiv
Salvato in:
| Autori principali: | Tang, Xiaoyu, Lin, Yixin, Dang, Ting, Zhang, Yuanfang, Cheng, Jintao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How Attention Shapes Emotion: A Comparative Study of Attention Mechanisms for Speech Emotion Recognition
di: Casals-Salvador, Marc, et al.
Pubblicazione: (2026)
di: Casals-Salvador, Marc, et al.
Pubblicazione: (2026)
Test-Time Adaptation for Speech Emotion Recognition
di: Dong, Jiaheng, et al.
Pubblicazione: (2026)
di: Dong, Jiaheng, et al.
Pubblicazione: (2026)
Multi-Scale Temporal Transformer For Speech Emotion Recognition
di: Li, Zhipeng, et al.
Pubblicazione: (2024)
di: Li, Zhipeng, et al.
Pubblicazione: (2024)
RawTFNet: A Lightweight CNN Architecture for Speech Anti-spoofing
di: Xiao, Yang, et al.
Pubblicazione: (2025)
di: Xiao, Yang, et al.
Pubblicazione: (2025)
EmoFormer: A Text-Independent Speech Emotion Recognition using a Hybrid Transformer-CNN model
di: Hasan, Rashedul, et al.
Pubblicazione: (2025)
di: Hasan, Rashedul, et al.
Pubblicazione: (2025)
Token-Level Logits Matter: A Closer Look at Speech Foundation Models for Ambiguous Emotion Recognition
di: Halim, Jule Valendo, et al.
Pubblicazione: (2025)
di: Halim, Jule Valendo, et al.
Pubblicazione: (2025)
Emo-bias: A Large Scale Evaluation of Social Bias on Speech Emotion Recognition
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention
di: Jiao, Xinxin, et al.
Pubblicazione: (2024)
di: Jiao, Xinxin, et al.
Pubblicazione: (2024)
Revisiting Modeling and Evaluation Approaches in Speech Emotion Recognition: Considering Subjectivity of Annotators and Ambiguity of Emotions
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2025)
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2025)
Searching for Effective Preprocessing Method and CNN-based Architecture with Efficient Channel Attention on Speech Emotion Recognition
di: Kim, Byunggun, et al.
Pubblicazione: (2024)
di: Kim, Byunggun, et al.
Pubblicazione: (2024)
Decoding Ambiguous Emotions with Test-Time Scaling in Audio-Language Models
di: Jia, Hong, et al.
Pubblicazione: (2026)
di: Jia, Hong, et al.
Pubblicazione: (2026)
Leveraging Content and Acoustic Representations for Speech Emotion Recognition
di: Dutta, Soumya, et al.
Pubblicazione: (2024)
di: Dutta, Soumya, et al.
Pubblicazione: (2024)
Leveraging Cross-Attention Transformer and Multi-Feature Fusion for Cross-Linguistic Speech Emotion Recognition
di: Zhao, Ruoyu, et al.
Pubblicazione: (2025)
di: Zhao, Ruoyu, et al.
Pubblicazione: (2025)
Speech Emotion Recognition with ASR Integration
di: Li, Yuanchao
Pubblicazione: (2026)
di: Li, Yuanchao
Pubblicazione: (2026)
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
di: Lin, Yuke, et al.
Pubblicazione: (2025)
di: Lin, Yuke, et al.
Pubblicazione: (2025)
EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
Reasoning Beyond Majority Vote: An Explainable SpeechLM Framework for Speech Emotion Recognition
di: Su, Bo-Hao, et al.
Pubblicazione: (2025)
di: Su, Bo-Hao, et al.
Pubblicazione: (2025)
Metadata-Enhanced Speech Emotion Recognition: Augmented Residual Integration and Co-Attention in Two-Stage Fine-Tuning
di: Wan, Zixiang, et al.
Pubblicazione: (2024)
di: Wan, Zixiang, et al.
Pubblicazione: (2024)
ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
di: Tang, Haobin, et al.
Pubblicazione: (2024)
di: Tang, Haobin, et al.
Pubblicazione: (2024)
Double Multi-Head Attention Multimodal System for Odyssey 2024 Speech Emotion Recognition Challenge
di: Costa, Federico, et al.
Pubblicazione: (2024)
di: Costa, Federico, et al.
Pubblicazione: (2024)
Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model
di: Ueda, Lucas, et al.
Pubblicazione: (2025)
di: Ueda, Lucas, et al.
Pubblicazione: (2025)
Identifying and Calibrating Overconfidence in Noisy Speech Recognition
di: Huo, Mingyue, et al.
Pubblicazione: (2025)
di: Huo, Mingyue, et al.
Pubblicazione: (2025)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
di: Shen, Siyuan, et al.
Pubblicazione: (2024)
di: Shen, Siyuan, et al.
Pubblicazione: (2024)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
di: Lin, Hsi-Che, et al.
Pubblicazione: (2024)
di: Lin, Hsi-Che, et al.
Pubblicazione: (2024)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
di: Zhao, Yan, et al.
Pubblicazione: (2024)
di: Zhao, Yan, et al.
Pubblicazione: (2024)
AdaLTM: Adaptive Layer-wise Task Vector Merging for Categorical Speech Emotion Recognition with ASR Knowledge Integration
di: Lee, Chia-Yu, et al.
Pubblicazione: (2026)
di: Lee, Chia-Yu, et al.
Pubblicazione: (2026)
HYFuse: Aligning Heterogeneous Speech Pre-Trained Representations in Hyperbolic Space for Speech Emotion Recognition
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
Re-Parameterization of Lightweight Transformer for On-Device Speech Emotion Recognition
di: Zhang, Zixing, et al.
Pubblicazione: (2024)
di: Zhang, Zixing, et al.
Pubblicazione: (2024)
EMO-Debias: Benchmarking Gender Debiasing Techniques in Multi-Label Speech Emotion Recognition
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language Model
di: Ihori, Mana, et al.
Pubblicazione: (2025)
di: Ihori, Mana, et al.
Pubblicazione: (2025)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
di: Hu, Cheng-Hung, et al.
Pubblicazione: (2025)
di: Hu, Cheng-Hung, et al.
Pubblicazione: (2025)
Attention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition
di: Yang, Qingran, et al.
Pubblicazione: (2026)
di: Yang, Qingran, et al.
Pubblicazione: (2026)
Temporal-Frequency State Space Duality: An Efficient Paradigm for Speech Emotion Recognition
di: Zhao, Jiaqi, et al.
Pubblicazione: (2024)
di: Zhao, Jiaqi, et al.
Pubblicazione: (2024)
Persian Speech Emotion Recognition by Fine-Tuning Transformers
di: Shayaninasab, Minoo, et al.
Pubblicazione: (2024)
di: Shayaninasab, Minoo, et al.
Pubblicazione: (2024)
B-GRPO: Unsupervised Speech Emotion Recognition based on Batched-Group Relative Policy Optimization
di: Gao, Yingying, et al.
Pubblicazione: (2026)
di: Gao, Yingying, et al.
Pubblicazione: (2026)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
di: Lin, Guan-Ting, et al.
Pubblicazione: (2024)
di: Lin, Guan-Ting, et al.
Pubblicazione: (2024)
Prosody as Supervision: Bridging the Non-Verbal--Verbal for Multilingual Speech Emotion Recognition
di: Girish, et al.
Pubblicazione: (2026)
di: Girish, et al.
Pubblicazione: (2026)
Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
di: Zhang, Wenda, et al.
Pubblicazione: (2026)
di: Zhang, Wenda, et al.
Pubblicazione: (2026)
Enhanced Speech Emotion Recognition with Efficient Channel Attention Guided Deep CNN-BiLSTM Framework
di: Kundu, Niloy Kumar, et al.
Pubblicazione: (2024)
di: Kundu, Niloy Kumar, et al.
Pubblicazione: (2024)
Dataset-Distillation Generative Model for Speech Emotion Recognition
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2024)
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2024)
Documenti analoghi
-
How Attention Shapes Emotion: A Comparative Study of Attention Mechanisms for Speech Emotion Recognition
di: Casals-Salvador, Marc, et al.
Pubblicazione: (2026) -
Test-Time Adaptation for Speech Emotion Recognition
di: Dong, Jiaheng, et al.
Pubblicazione: (2026) -
Multi-Scale Temporal Transformer For Speech Emotion Recognition
di: Li, Zhipeng, et al.
Pubblicazione: (2024) -
RawTFNet: A Lightweight CNN Architecture for Speech Anti-spoofing
di: Xiao, Yang, et al.
Pubblicazione: (2025) -
EmoFormer: A Text-Independent Speech Emotion Recognition using a Hybrid Transformer-CNN model
di: Hasan, Rashedul, et al.
Pubblicazione: (2025)