Speech Swin-Transformer: Exploring a Hierarchical Transformer with Shifted Windows for Speech Emotion Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yong, Lu, Cheng, Lian, Hailun, Zhao, Yan, Schuller, Björn, Zong, Yuan, Zheng, Wenming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving Speaker-independent Speech Emotion Recognition Using Dynamic Joint Distribution Adaptation
von: Lu, Cheng, et al.
Veröffentlicht: (2024)
von: Lu, Cheng, et al.
Veröffentlicht: (2024)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
INTERSPEECH 2009 Emotion Challenge Revisited: Benchmarking 15 Years of Progress in Speech Emotion Recognition
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024)
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024)
Explainable Speech Emotion Recognition: Weighted Attribute Fairness to Model Demographic Contributions to Social Bias
von: Ogunnubi, Tomisin, et al.
Veröffentlicht: (2026)
von: Ogunnubi, Tomisin, et al.
Veröffentlicht: (2026)
Temporal Label Hierachical Network for Compound Emotion Recognition
von: Li, Sunan, et al.
Veröffentlicht: (2024)
von: Li, Sunan, et al.
Veröffentlicht: (2024)
Expressivity and Speech Synthesis
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024)
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024)
PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion
von: Qi, Tianhua, et al.
Veröffentlicht: (2024)
von: Qi, Tianhua, et al.
Veröffentlicht: (2024)
Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
Explainable Transformer-CNN Fusion for Noise-Robust Speech Emotion Recognition
von: Chakrabarty, Sudip, et al.
Veröffentlicht: (2025)
von: Chakrabarty, Sudip, et al.
Veröffentlicht: (2025)
Hybrid CNN-Transformer Architecture for Arabic Speech Emotion Recognition
von: Gheffari, Youcef Soufiane, et al.
Veröffentlicht: (2026)
von: Gheffari, Youcef Soufiane, et al.
Veröffentlicht: (2026)
Re-Parameterization of Lightweight Transformer for On-Device Speech Emotion Recognition
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
Leveraging Cross-Attention Transformer and Multi-Feature Fusion for Cross-Linguistic Speech Emotion Recognition
von: Zhao, Ruoyu, et al.
Veröffentlicht: (2025)
von: Zhao, Ruoyu, et al.
Veröffentlicht: (2025)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
Modeling Emotional Trajectories in Written Stories Utilizing Transformers and Weakly-Supervised Learning
von: Christ, Lukas, et al.
Veröffentlicht: (2024)
von: Christ, Lukas, et al.
Veröffentlicht: (2024)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
Speech Recognition Transformers: Topological-lingualism Perspective
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
On the Emotion Understanding of Synthesized Speech
von: Ge, Yuan, et al.
Veröffentlicht: (2026)
von: Ge, Yuan, et al.
Veröffentlicht: (2026)
HAFFormer: A Hierarchical Attention-Free Framework for Alzheimer's Disease Detection From Spontaneous Speech
von: Dong, Zhongren, et al.
Veröffentlicht: (2024)
von: Dong, Zhongren, et al.
Veröffentlicht: (2024)
CMATH: Cross-Modality Augmented Transformer with Hierarchical Variational Distillation for Multimodal Emotion Recognition in Conversation
von: Zhu, Xiaofei, et al.
Veröffentlicht: (2024)
von: Zhu, Xiaofei, et al.
Veröffentlicht: (2024)
Feature-Augmented Transformers for Robust AI-Text Detection Across Domains and Generators
von: Mady, Mohamed, et al.
Veröffentlicht: (2026)
von: Mady, Mohamed, et al.
Veröffentlicht: (2026)
Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition
von: Sun, Haiyang, et al.
Veröffentlicht: (2023)
von: Sun, Haiyang, et al.
Veröffentlicht: (2023)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
Emotion and Intent Joint Understanding in Multimodal Conversation: A Benchmarking Dataset
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis
von: He, Xiangheng, et al.
Veröffentlicht: (2024)
von: He, Xiangheng, et al.
Veröffentlicht: (2024)
Unrequited Emotions: Investigating the Gaps in Motivation and Practice in Speech Emotion Recognition Research
von: Wong, Taryn, et al.
Veröffentlicht: (2026)
von: Wong, Taryn, et al.
Veröffentlicht: (2026)
Transfer Learning of Transformer-based Speech Recognition Models from Czech to Slovak
von: Lehečka, Jan, et al.
Veröffentlicht: (2023)
von: Lehečka, Jan, et al.
Veröffentlicht: (2023)
Automatic Speech Recognition with BERT and CTC Transformers: A Review
von: Djeffal, Noussaiba, et al.
Veröffentlicht: (2024)
von: Djeffal, Noussaiba, et al.
Veröffentlicht: (2024)
Semantic Differentiation in Speech Emotion Recognition: Insights from Descriptive and Expressive Speech Roles
von: Guo, Rongchen, et al.
Veröffentlicht: (2025)
von: Guo, Rongchen, et al.
Veröffentlicht: (2025)
On the Contribution of Lexical Features to Speech Emotion Recognition
von: Combei, David
Veröffentlicht: (2025)
von: Combei, David
Veröffentlicht: (2025)
Enhancing DR Classification with Swin Transformer and Shifted Window Attention
von: Boulaabi, Meher, et al.
Veröffentlicht: (2025)
von: Boulaabi, Meher, et al.
Veröffentlicht: (2025)
DANCER: Entity Description Augmented Named Entity Corrector for Automatic Speech Recognition
von: Wang, Yi-Cheng, et al.
Veröffentlicht: (2024)
von: Wang, Yi-Cheng, et al.
Veröffentlicht: (2024)
Evaluating Emotion Recognition in Spoken Language Models on Emotionally Incongruent Speech
von: Corrêa, Pedro, et al.
Veröffentlicht: (2025)
von: Corrêa, Pedro, et al.
Veröffentlicht: (2025)
Investigating the Impact of Word Informativeness on Speech Emotion Recognition
von: Kakouros, Sofoklis
Veröffentlicht: (2025)
von: Kakouros, Sofoklis
Veröffentlicht: (2025)
Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
von: Zhu, Yongxin, et al.
Veröffentlicht: (2024)
von: Zhu, Yongxin, et al.
Veröffentlicht: (2024)
EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses
von: Xu, Shuhao, et al.
Veröffentlicht: (2026)
von: Xu, Shuhao, et al.
Veröffentlicht: (2026)
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025
von: Ferreira, Alef Iury Siqueira, et al.
Veröffentlicht: (2025)
von: Ferreira, Alef Iury Siqueira, et al.
Veröffentlicht: (2025)
EMO-Debias: Benchmarking Gender Debiasing Techniques in Multi-Label Speech Emotion Recognition
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Improving Speaker-independent Speech Emotion Recognition Using Dynamic Joint Distribution Adaptation
von: Lu, Cheng, et al.
Veröffentlicht: (2024) -
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
von: Zhao, Yan, et al.
Veröffentlicht: (2024) -
INTERSPEECH 2009 Emotion Challenge Revisited: Benchmarking 15 Years of Progress in Speech Emotion Recognition
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024) -
Explainable Speech Emotion Recognition: Weighted Attribute Fairness to Model Demographic Contributions to Social Bias
von: Ogunnubi, Tomisin, et al.
Veröffentlicht: (2026) -
Temporal Label Hierachical Network for Compound Emotion Recognition
von: Li, Sunan, et al.
Veröffentlicht: (2024)