SpectroFusion-ViT: A Lightweight Transformer for Speech Emotion Recognition Using Harmonic Mel-Chroma Fusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ahmed, Faria, Chowdhury, Rafi Hassan, Moon, Fatema Tuz Zohora, Ahmed, Sabbir |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MangoLeafViT: Leveraging Lightweight Vision Transformer with Runtime Augmentation for Efficient Mango Leaf Disease Classification
von: Chowdhury, Rafi Hassan, et al.
Veröffentlicht: (2025)
von: Chowdhury, Rafi Hassan, et al.
Veröffentlicht: (2025)
Explainable Transformer-CNN Fusion for Noise-Robust Speech Emotion Recognition
von: Chakrabarty, Sudip, et al.
Veröffentlicht: (2025)
von: Chakrabarty, Sudip, et al.
Veröffentlicht: (2025)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
Enhancing Speech Emotion Recognition with Multi-Task Learning and Dynamic Feature Fusion
von: Wang, Honghong, et al.
Veröffentlicht: (2025)
von: Wang, Honghong, et al.
Veröffentlicht: (2025)
Deep Learning for Speech Emotion Recognition: A CNN Approach Utilizing Mel Spectrograms
von: Penumajji, Niketa
Veröffentlicht: (2025)
von: Penumajji, Niketa
Veröffentlicht: (2025)
Real-Time Speech Enhancement via a Hybrid ViT: A Dual-Input Acoustic-Image Feature Fusion
von: Bahmei, Behnaz, et al.
Veröffentlicht: (2025)
von: Bahmei, Behnaz, et al.
Veröffentlicht: (2025)
Re-Parameterization of Lightweight Transformer for On-Device Speech Emotion Recognition
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
LPGNet: A Lightweight Network with Parallel Attention and Gated Fusion for Multimodal Emotion Recognition
von: He, Zhining, et al.
Veröffentlicht: (2025)
von: He, Zhining, et al.
Veröffentlicht: (2025)
Leveraging Cross-Attention Transformer and Multi-Feature Fusion for Cross-Linguistic Speech Emotion Recognition
von: Zhao, Ruoyu, et al.
Veröffentlicht: (2025)
von: Zhao, Ruoyu, et al.
Veröffentlicht: (2025)
DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
Bimodal Connection Attention Fusion for Speech Emotion Recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025
von: Ferreira, Alef Iury Siqueira, et al.
Veröffentlicht: (2025)
von: Ferreira, Alef Iury Siqueira, et al.
Veröffentlicht: (2025)
Emotion Detection in Speech Using Lightweight and Transformer-Based Models: A Comparative and Ablation Study
von: Onyekwelu-Udoka, Lucky, et al.
Veröffentlicht: (2025)
von: Onyekwelu-Udoka, Lucky, et al.
Veröffentlicht: (2025)
MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition
von: Sun, Haiyang, et al.
Veröffentlicht: (2023)
von: Sun, Haiyang, et al.
Veröffentlicht: (2023)
MelShield: Robust Mel-Domain Audio Watermarking for Provenance Attribution of AI Generated Synthesized Speech
von: Jin, Yutong, et al.
Veröffentlicht: (2026)
von: Jin, Yutong, et al.
Veröffentlicht: (2026)
CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR
von: Shao, Nian, et al.
Veröffentlicht: (2025)
von: Shao, Nian, et al.
Veröffentlicht: (2025)
Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech Recognition
von: Liu, Lei, et al.
Veröffentlicht: (2024)
von: Liu, Lei, et al.
Veröffentlicht: (2024)
Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement
von: Yang, Yujie, et al.
Veröffentlicht: (2025)
von: Yang, Yujie, et al.
Veröffentlicht: (2025)
Multi-Channel Speech Enhancement for Cocktail Party Speech Emotion Recognition
von: Chen, Youjun, et al.
Veröffentlicht: (2026)
von: Chen, Youjun, et al.
Veröffentlicht: (2026)
Persian Speech Emotion Recognition by Fine-Tuning Transformers
von: Shayaninasab, Minoo, et al.
Veröffentlicht: (2024)
von: Shayaninasab, Minoo, et al.
Veröffentlicht: (2024)
Graph Embedding with Mel-spectrograms for Underwater Acoustic Target Recognition
von: Feng, Sheng, et al.
Veröffentlicht: (2025)
von: Feng, Sheng, et al.
Veröffentlicht: (2025)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
Recovering Performance in Speech Emotion Recognition from Discrete Tokens via Multi-Layer Fusion and Paralinguistic Feature Integration
von: Sun, Esther, et al.
Veröffentlicht: (2026)
von: Sun, Esther, et al.
Veröffentlicht: (2026)
Spectro-Temporal Modulation Representation Framework for Human-Imitated Speech Detection
von: Zaman, Khalid, et al.
Veröffentlicht: (2026)
von: Zaman, Khalid, et al.
Veröffentlicht: (2026)
Hybrid CNN-Transformer Architecture for Arabic Speech Emotion Recognition
von: Gheffari, Youcef Soufiane, et al.
Veröffentlicht: (2026)
von: Gheffari, Youcef Soufiane, et al.
Veröffentlicht: (2026)
MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction
von: He, Jiajun, et al.
Veröffentlicht: (2024)
von: He, Jiajun, et al.
Veröffentlicht: (2024)
WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition
von: Li, Feng, et al.
Veröffentlicht: (2024)
von: Li, Feng, et al.
Veröffentlicht: (2024)
Efficient Finetuning for Dimensional Speech Emotion Recognition in the Age of Transformers
von: Sampath, Aneesha, et al.
Veröffentlicht: (2025)
von: Sampath, Aneesha, et al.
Veröffentlicht: (2025)
Emotion Recognition in Multi-Speaker Conversations through Speaker Identification, Knowledge Distillation, and Hierarchical Fusion
von: Li, Xiao, et al.
Veröffentlicht: (2025)
von: Li, Xiao, et al.
Veröffentlicht: (2025)
Speech Emotion Recognition with ASR Integration
von: Li, Yuanchao
Veröffentlicht: (2026)
von: Li, Yuanchao
Veröffentlicht: (2026)
Speech Representation Analysis based on Inter- and Intra-Model Similarities
von: Kheir, Yassine El, et al.
Veröffentlicht: (2024)
von: Kheir, Yassine El, et al.
Veröffentlicht: (2024)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
von: Li, Qifei, et al.
Veröffentlicht: (2024)
von: Li, Qifei, et al.
Veröffentlicht: (2024)
dMel: Speech Tokenization made Simple
von: Bai, Richard He, et al.
Veröffentlicht: (2024)
von: Bai, Richard He, et al.
Veröffentlicht: (2024)
A Novel Fusion Architecture for PD Detection Using Semi-Supervised Speech Embeddings
von: Adnan, Tariq, et al.
Veröffentlicht: (2024)
von: Adnan, Tariq, et al.
Veröffentlicht: (2024)
Cross-Cultural Bias in Mel-Scale Representations: Evidence and Alternatives from Speech and Music
von: Chauhan, Shivam, et al.
Veröffentlicht: (2026)
von: Chauhan, Shivam, et al.
Veröffentlicht: (2026)
Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
von: Aboeitta, Ahmed, et al.
Veröffentlicht: (2025)
von: Aboeitta, Ahmed, et al.
Veröffentlicht: (2025)
An LSTM-Based Chord Generation System Using Chroma Histogram Representations
von: Hardwick, Jack
Veröffentlicht: (2024)
von: Hardwick, Jack
Veröffentlicht: (2024)
Ähnliche Einträge
-
MangoLeafViT: Leveraging Lightweight Vision Transformer with Runtime Augmentation for Efficient Mango Leaf Disease Classification
von: Chowdhury, Rafi Hassan, et al.
Veröffentlicht: (2025) -
Explainable Transformer-CNN Fusion for Noise-Robust Speech Emotion Recognition
von: Chakrabarty, Sudip, et al.
Veröffentlicht: (2025) -
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023) -
Enhancing Speech Emotion Recognition with Multi-Task Learning and Dynamic Feature Fusion
von: Wang, Honghong, et al.
Veröffentlicht: (2025) -
Deep Learning for Speech Emotion Recognition: A CNN Approach Utilizing Mel Spectrograms
von: Penumajji, Niketa
Veröffentlicht: (2025)