Gespeichert in:
| Hauptverfasser: | He, Zhining, Xiao, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2508.08925 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Double Multi-Head Attention Multimodal System for Odyssey 2024 Speech Emotion Recognition Challenge
von: Costa, Federico, et al.
Veröffentlicht: (2024)
von: Costa, Federico, et al.
Veröffentlicht: (2024)
Emotion Recognition in Multi-Speaker Conversations through Speaker Identification, Knowledge Distillation, and Hierarchical Fusion
von: Li, Xiao, et al.
Veröffentlicht: (2025)
von: Li, Xiao, et al.
Veröffentlicht: (2025)
Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion
von: Deng, Zeyu, et al.
Veröffentlicht: (2025)
von: Deng, Zeyu, et al.
Veröffentlicht: (2025)
PARROT: Synergizing Mamba and Attention-based SSL Pre-Trained Models via Parallel Branch Hadamard Optimal Transport for Speech Emotion Recognition
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Recursive Joint Cross-Modal Attention for Multimodal Fusion in Dimensional Emotion Recognition
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2024)
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2024)
A Survey on Multimodal Music Emotion Recognition
von: Liyanarachchi, Rashini, et al.
Veröffentlicht: (2025)
von: Liyanarachchi, Rashini, et al.
Veröffentlicht: (2025)
LLM supervised Pre-training for Multimodal Emotion Recognition in Conversations
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
Low-Complexity Acoustic Scene Classification Using Parallel Attention-Convolution Network
von: Li, Yanxiong, et al.
Veröffentlicht: (2024)
von: Li, Yanxiong, et al.
Veröffentlicht: (2024)
Attention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition
von: Yang, Qingran, et al.
Veröffentlicht: (2026)
von: Yang, Qingran, et al.
Veröffentlicht: (2026)
Leveraging Cross-Attention Transformer and Multi-Feature Fusion for Cross-Linguistic Speech Emotion Recognition
von: Zhao, Ruoyu, et al.
Veröffentlicht: (2025)
von: Zhao, Ruoyu, et al.
Veröffentlicht: (2025)
MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model
von: Ueda, Lucas, et al.
Veröffentlicht: (2025)
von: Ueda, Lucas, et al.
Veröffentlicht: (2025)
TelME: Teacher-leading Multimodal Fusion Network for Emotion Recognition in Conversation
von: Yun, Taeyang, et al.
Veröffentlicht: (2024)
von: Yun, Taeyang, et al.
Veröffentlicht: (2024)
Re-Parameterization of Lightweight Transformer for On-Device Speech Emotion Recognition
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
Multimodal Emotion Recognition from Raw Audio with Sinc-convolution
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
EfficientASR: Speech Recognition Network Compression via Attention Redundancy and Chunk-Level FFN Optimization
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
Bridging Speech Emotion Recognition and Personality: Dataset and Temporal Interaction Condition Network
von: Gao, Yuan, et al.
Veröffentlicht: (2025)
von: Gao, Yuan, et al.
Veröffentlicht: (2025)
TBDM-Net: Bidirectional Dense Networks with Gender Information for Speech Emotion Recognition
von: Striletchi, Vlad, et al.
Veröffentlicht: (2024)
von: Striletchi, Vlad, et al.
Veröffentlicht: (2024)
Metadata-Enhanced Speech Emotion Recognition: Augmented Residual Integration and Co-Attention in Two-Stage Fine-Tuning
von: Wan, Zixiang, et al.
Veröffentlicht: (2024)
von: Wan, Zixiang, et al.
Veröffentlicht: (2024)
PERSONA: An Application for Emotion Recognition, Gender Recognition and Age Estimation
von: Koshal, Devyani, et al.
Veröffentlicht: (2024)
von: Koshal, Devyani, et al.
Veröffentlicht: (2024)
EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognition
von: Li, Pengcheng, et al.
Veröffentlicht: (2025)
von: Li, Pengcheng, et al.
Veröffentlicht: (2025)
Speech Emotion Recognition with ASR Integration
von: Li, Yuanchao
Veröffentlicht: (2026)
von: Li, Yuanchao
Veröffentlicht: (2026)
Speaker Recognition Using Isomorphic Graph Attention Network Based Pooling on Self-Supervised Representation
von: Ge, Zirui, et al.
Veröffentlicht: (2023)
von: Ge, Zirui, et al.
Veröffentlicht: (2023)
MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction
von: He, Jiajun, et al.
Veröffentlicht: (2024)
von: He, Jiajun, et al.
Veröffentlicht: (2024)
Textless and Non-Parallel Speech-to-Speech Emotion Style Transfer
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
ArabEmoNet: A Lightweight Hybrid 2D CNN-BiLSTM Model with Attention for Robust Arabic Speech Emotion Recognition
von: Abouzeid, Ali, et al.
Veröffentlicht: (2025)
von: Abouzeid, Ali, et al.
Veröffentlicht: (2025)
Temporal-Frequency State Space Duality: An Efficient Paradigm for Speech Emotion Recognition
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2024)
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2024)
Reverse Attention for Lightweight Speech Enhancement on Edge Devices
von: Ojha, Shuubham, et al.
Veröffentlicht: (2025)
von: Ojha, Shuubham, et al.
Veröffentlicht: (2025)
Bimodal Connection Attention Fusion for Speech Emotion Recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
RawTFNet: A Lightweight CNN Architecture for Speech Anti-spoofing
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition
von: Sun, Haiyang, et al.
Veröffentlicht: (2023)
von: Sun, Haiyang, et al.
Veröffentlicht: (2023)
Audio-Guided Fusion Techniques for Multimodal Emotion Analysis
von: Shi, Pujin, et al.
Veröffentlicht: (2024)
von: Shi, Pujin, et al.
Veröffentlicht: (2024)
Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
Leveraging Label Potential for Enhanced Multimodal Emotion Recognition
von: Shao, Xuechun, et al.
Veröffentlicht: (2025)
von: Shao, Xuechun, et al.
Veröffentlicht: (2025)
A Lightweight Slot-Attention Framework for Multi-Instrument Multi-Pitch Estimation
von: Taenzer, Michael
Veröffentlicht: (2026)
von: Taenzer, Michael
Veröffentlicht: (2026)
MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
THAI Speech Emotion Recognition (THAI-SER) corpus
von: Wongpithayadisai, Jilamika, et al.
Veröffentlicht: (2025)
von: Wongpithayadisai, Jilamika, et al.
Veröffentlicht: (2025)
TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
von: Yang, Cheng-Yeh, et al.
Veröffentlicht: (2026)
von: Yang, Cheng-Yeh, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Double Multi-Head Attention Multimodal System for Odyssey 2024 Speech Emotion Recognition Challenge
von: Costa, Federico, et al.
Veröffentlicht: (2024) -
Emotion Recognition in Multi-Speaker Conversations through Speaker Identification, Knowledge Distillation, and Hierarchical Fusion
von: Li, Xiao, et al.
Veröffentlicht: (2025) -
Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion
von: Deng, Zeyu, et al.
Veröffentlicht: (2025) -
PARROT: Synergizing Mamba and Attention-based SSL Pre-Trained Models via Parallel Branch Hadamard Optimal Transport for Speech Emotion Recognition
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025) -
Recursive Joint Cross-Modal Attention for Multimodal Fusion in Dimensional Emotion Recognition
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2024)