BSC-UPC at EmoSPeech-IberLEF2024: Attention Pooling for Emotion Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Casals-Salvador, Marc, Costa, Federico, India, Miquel, Hernando, Javier |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Double Multi-Head Attention Multimodal System for Odyssey 2024 Speech Emotion Recognition Challenge
por: Costa, Federico, et al.
Publicado: (2024)
por: Costa, Federico, et al.
Publicado: (2024)
How Attention Shapes Emotion: A Comparative Study of Attention Mechanisms for Speech Emotion Recognition
por: Casals-Salvador, Marc, et al.
Publicado: (2026)
por: Casals-Salvador, Marc, et al.
Publicado: (2026)
Speaker Characterization by means of Attention Pooling
por: Costa, Federico, et al.
Publicado: (2024)
por: Costa, Federico, et al.
Publicado: (2024)
Emo-bias: A Large Scale Evaluation of Social Bias on Speech Emotion Recognition
por: Lin, Yi-Cheng, et al.
Publicado: (2024)
por: Lin, Yi-Cheng, et al.
Publicado: (2024)
Language Modelling for Speaker Diarization in Telephonic Interviews
por: India, Miquel, et al.
Publicado: (2025)
por: India, Miquel, et al.
Publicado: (2025)
EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition
por: Thimonier, Hugo, et al.
Publicado: (2025)
por: Thimonier, Hugo, et al.
Publicado: (2025)
EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
por: Yang, Yiqing, et al.
Publicado: (2025)
por: Yang, Yiqing, et al.
Publicado: (2025)
KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis
por: Abilbekov, Adal, et al.
Publicado: (2024)
por: Abilbekov, Adal, et al.
Publicado: (2024)
EmoFormer: A Text-Independent Speech Emotion Recognition using a Hybrid Transformer-CNN model
por: Hasan, Rashedul, et al.
Publicado: (2025)
por: Hasan, Rashedul, et al.
Publicado: (2025)
EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech Tokens
por: Park, Joonyong, et al.
Publicado: (2025)
por: Park, Joonyong, et al.
Publicado: (2025)
EmoFake: An Initial Dataset for Emotion Fake Audio Detection
por: Zhao, Yan, et al.
Publicado: (2022)
por: Zhao, Yan, et al.
Publicado: (2022)
Speech Emotion Recognition Via CNN-Transformer and Multidimensional Attention Mechanism
por: Tang, Xiaoyu, et al.
Publicado: (2024)
por: Tang, Xiaoyu, et al.
Publicado: (2024)
EmoOmni: Bridging Emotional Understanding and Expression in Omni-Modal LLMs
por: Tian, Wenjie, et al.
Publicado: (2026)
por: Tian, Wenjie, et al.
Publicado: (2026)
Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model
por: Ueda, Lucas, et al.
Publicado: (2025)
por: Ueda, Lucas, et al.
Publicado: (2025)
EmoTech: A Multi-modal Speech Emotion Recognition Using Multi-source Low-level Information with Hybrid Recurrent Network
por: Avro, Shamin Bin Habib, et al.
Publicado: (2025)
por: Avro, Shamin Bin Habib, et al.
Publicado: (2025)
Speaker Recognition Using Isomorphic Graph Attention Network Based Pooling on Self-Supervised Representation
por: Ge, Zirui, et al.
Publicado: (2023)
por: Ge, Zirui, et al.
Publicado: (2023)
ArabEmoNet: A Lightweight Hybrid 2D CNN-BiLSTM Model with Attention for Robust Arabic Speech Emotion Recognition
por: Abouzeid, Ali, et al.
Publicado: (2025)
por: Abouzeid, Ali, et al.
Publicado: (2025)
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
por: Cho, Deok-Hyeon, et al.
Publicado: (2025)
por: Cho, Deok-Hyeon, et al.
Publicado: (2025)
Infant Cry Emotion Recognition Using Improved ECAPA-TDNN with Multiscale Feature Fusion and Attention Enhancement
por: Zhou, Junyu, et al.
Publicado: (2025)
por: Zhou, Junyu, et al.
Publicado: (2025)
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
por: Gao, Xiaoxue, et al.
Publicado: (2024)
por: Gao, Xiaoxue, et al.
Publicado: (2024)
EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
por: Cong, Gaoxiang, et al.
Publicado: (2024)
por: Cong, Gaoxiang, et al.
Publicado: (2024)
On the Use of Audio to Improve Dialogue Policies
por: Roncel, Daniel, et al.
Publicado: (2024)
por: Roncel, Daniel, et al.
Publicado: (2024)
Quantifying Cross-Lingual Transfer in Paralinguistic Speech Tasks
por: Buitrago, Pol, et al.
Publicado: (2026)
por: Buitrago, Pol, et al.
Publicado: (2026)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
por: Zhang, Hezhao, et al.
Publicado: (2026)
por: Zhang, Hezhao, et al.
Publicado: (2026)
DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches
por: Vu, Tai
Publicado: (2025)
por: Vu, Tai
Publicado: (2025)
LPGNet: A Lightweight Network with Parallel Attention and Gated Fusion for Multimodal Emotion Recognition
por: He, Zhining, et al.
Publicado: (2025)
por: He, Zhining, et al.
Publicado: (2025)
Multi-Scale Temporal Transformer For Speech Emotion Recognition
por: Li, Zhipeng, et al.
Publicado: (2024)
por: Li, Zhipeng, et al.
Publicado: (2024)
Leveraging Content and Acoustic Representations for Speech Emotion Recognition
por: Dutta, Soumya, et al.
Publicado: (2024)
por: Dutta, Soumya, et al.
Publicado: (2024)
Temporal Attention Pooling for Frequency Dynamic Convolution in Sound Event Detection
por: Nam, Hyeonuk, et al.
Publicado: (2025)
por: Nam, Hyeonuk, et al.
Publicado: (2025)
Revisiting Modeling and Evaluation Approaches in Speech Emotion Recognition: Considering Subjectivity of Annotators and Ambiguity of Emotions
por: Chou, Huang-Cheng, et al.
Publicado: (2025)
por: Chou, Huang-Cheng, et al.
Publicado: (2025)
Enhancing Speech Emotion Recognition through Segmental Average Pooling of Self-Supervised Learning Features
por: Hyeon, Jonghwan, et al.
Publicado: (2024)
por: Hyeon, Jonghwan, et al.
Publicado: (2024)
Evaluating Emotion Recognition in Spoken Language Models on Emotionally Incongruent Speech
por: Corrêa, Pedro, et al.
Publicado: (2025)
por: Corrêa, Pedro, et al.
Publicado: (2025)
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
por: Cho, Deok-Hyeon, et al.
Publicado: (2024)
por: Cho, Deok-Hyeon, et al.
Publicado: (2024)
SoCov: Semi-Orthogonal Parametric Pooling of Covariance Matrix for Speaker Recognition
por: Li, Rongjin, et al.
Publicado: (2025)
por: Li, Rongjin, et al.
Publicado: (2025)
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
por: Cho, Deok-Hyeon, et al.
Publicado: (2024)
por: Cho, Deok-Hyeon, et al.
Publicado: (2024)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
por: Shen, Siyuan, et al.
Publicado: (2024)
por: Shen, Siyuan, et al.
Publicado: (2024)
EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
por: Bian, Weizhen, et al.
Publicado: (2024)
por: Bian, Weizhen, et al.
Publicado: (2024)
Metadata-Enhanced Speech Emotion Recognition: Augmented Residual Integration and Co-Attention in Two-Stage Fine-Tuning
por: Wan, Zixiang, et al.
Publicado: (2024)
por: Wan, Zixiang, et al.
Publicado: (2024)
Speech Emotion Recognition Using Fine-Tuned DWFormer:A Study on Track 1 of the IERPChallenge 2024
por: Wang, Honghong, et al.
Publicado: (2025)
por: Wang, Honghong, et al.
Publicado: (2025)
EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
por: Ma, Ziyang, et al.
Publicado: (2024)
por: Ma, Ziyang, et al.
Publicado: (2024)
Ejemplares similares
-
Double Multi-Head Attention Multimodal System for Odyssey 2024 Speech Emotion Recognition Challenge
por: Costa, Federico, et al.
Publicado: (2024) -
How Attention Shapes Emotion: A Comparative Study of Attention Mechanisms for Speech Emotion Recognition
por: Casals-Salvador, Marc, et al.
Publicado: (2026) -
Speaker Characterization by means of Attention Pooling
por: Costa, Federico, et al.
Publicado: (2024) -
Emo-bias: A Large Scale Evaluation of Social Bias on Speech Emotion Recognition
por: Lin, Yi-Cheng, et al.
Publicado: (2024) -
Language Modelling for Speaker Diarization in Telephonic Interviews
por: India, Miquel, et al.
Publicado: (2025)