Emotion Detection in Speech Using Lightweight and Transformer-Based Models: A Comparative and Ablation Study
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Onyekwelu-Udoka, Lucky, Islam, Md Shafiqul, Hasan, Md Shahedul |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SS-DPPN: A self-supervised dual-path foundation model for the generalizable cardiac audio representation
von: Muna, Ummy Maria, et al.
Veröffentlicht: (2025)
von: Muna, Ummy Maria, et al.
Veröffentlicht: (2025)
Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla
von: Ridoy, Md Sazzadul Islam, et al.
Veröffentlicht: (2025)
von: Ridoy, Md Sazzadul Islam, et al.
Veröffentlicht: (2025)
Zero-Shot to Zero-Lies: Detecting Bengali Deepfake Audio through Transfer Learning
von: Samu, Most. Sharmin Sultana, et al.
Veröffentlicht: (2025)
von: Samu, Most. Sharmin Sultana, et al.
Veröffentlicht: (2025)
An Investigation Into Various Approaches For Bengali Long-Form Speech Transcription and Bengali Speaker Diarization
von: Jahan, Epshita, et al.
Veröffentlicht: (2026)
von: Jahan, Epshita, et al.
Veröffentlicht: (2026)
Explainable Multi-Modal Deep Learning for Automatic Detection of Lung Diseases from Respiratory Audio Signals
von: Saky, S M Asiful Islam, et al.
Veröffentlicht: (2025)
von: Saky, S M Asiful Islam, et al.
Veröffentlicht: (2025)
MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages
von: Sailor, Hardik B., et al.
Veröffentlicht: (2025)
von: Sailor, Hardik B., et al.
Veröffentlicht: (2025)
A Novel Hybrid Deep Learning Technique for Speech Emotion Detection using Feature Engineering
von: Chowdhury, Shahana Yasmin, et al.
Veröffentlicht: (2025)
von: Chowdhury, Shahana Yasmin, et al.
Veröffentlicht: (2025)
SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
von: Nayeem, Md., et al.
Veröffentlicht: (2025)
von: Nayeem, Md., et al.
Veröffentlicht: (2025)
Speech Emotion Recognition Using MFCC Features and LSTM-Based Deep Learning Model
von: Oluwademilade, Adelekun, et al.
Veröffentlicht: (2026)
von: Oluwademilade, Adelekun, et al.
Veröffentlicht: (2026)
Hybrid CNN-Transformer Architecture for Arabic Speech Emotion Recognition
von: Gheffari, Youcef Soufiane, et al.
Veröffentlicht: (2026)
von: Gheffari, Youcef Soufiane, et al.
Veröffentlicht: (2026)
LSZone: A Lightweight Spatial Information Modeling Architecture for Real-time In-car Multi-zone Speech Separation
von: Chen, Jun, et al.
Veröffentlicht: (2025)
von: Chen, Jun, et al.
Veröffentlicht: (2025)
Persian Speech Emotion Recognition by Fine-Tuning Transformers
von: Shayaninasab, Minoo, et al.
Veröffentlicht: (2024)
von: Shayaninasab, Minoo, et al.
Veröffentlicht: (2024)
A Lightweight Pipeline for Noisy Speech Voice Cloning and Accurate Lip Sync Synthesis
von: Amir, Javeria, et al.
Veröffentlicht: (2025)
von: Amir, Javeria, et al.
Veröffentlicht: (2025)
When Denoising Hinders: Revisiting Zero-Shot ASR with SAM-Audio and Whisper
von: Islam, Akif, et al.
Veröffentlicht: (2026)
von: Islam, Akif, et al.
Veröffentlicht: (2026)
A General Model for Deepfake Speech Detection: Diverse Bonafide Resources or Diverse AI-Based Generators
von: Pham, Lam, et al.
Veröffentlicht: (2026)
von: Pham, Lam, et al.
Veröffentlicht: (2026)
EMO-TTA: Improving Test-Time Adaptation of Audio-Language Models for Speech Emotion Recognition
von: Shi, Jiacheng, et al.
Veröffentlicht: (2025)
von: Shi, Jiacheng, et al.
Veröffentlicht: (2025)
Speech Emotion Recognition via Entropy-Aware Score Selection
von: Chua, ChenYi, et al.
Veröffentlicht: (2025)
von: Chua, ChenYi, et al.
Veröffentlicht: (2025)
State Space Models for Bioacoustics: A Comparative Evaluation with Transformers
von: Tang, Chengyu, et al.
Veröffentlicht: (2025)
von: Tang, Chengyu, et al.
Veröffentlicht: (2025)
Efficient Finetuning for Dimensional Speech Emotion Recognition in the Age of Transformers
von: Sampath, Aneesha, et al.
Veröffentlicht: (2025)
von: Sampath, Aneesha, et al.
Veröffentlicht: (2025)
DM-Codec: Distilling Multimodal Representations for Speech Tokenization
von: Ahasan, Md Mubtasim, et al.
Veröffentlicht: (2024)
von: Ahasan, Md Mubtasim, et al.
Veröffentlicht: (2024)
Deep Learning for Speech Emotion Recognition: A CNN Approach Utilizing Mel Spectrograms
von: Penumajji, Niketa
Veröffentlicht: (2025)
von: Penumajji, Niketa
Veröffentlicht: (2025)
Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement
von: Yang, Jianing, et al.
Veröffentlicht: (2025)
von: Yang, Jianing, et al.
Veröffentlicht: (2025)
EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis
von: Zhou, Li, et al.
Veröffentlicht: (2026)
von: Zhou, Li, et al.
Veröffentlicht: (2026)
Quantum Kernels for Audio Deepfake Detection Using Spectrogram Patch Features
von: Amin, Lisan Al, et al.
Veröffentlicht: (2026)
von: Amin, Lisan Al, et al.
Veröffentlicht: (2026)
SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
Whisper Speaker Identification: Leveraging Pre-Trained Multilingual Transformers for Robust Speaker Embeddings
von: Emon, Jakaria Islam, et al.
Veröffentlicht: (2025)
von: Emon, Jakaria Islam, et al.
Veröffentlicht: (2025)
Enhanced Speech Emotion Recognition with Efficient Channel Attention Guided Deep CNN-BiLSTM Framework
von: Kundu, Niloy Kumar, et al.
Veröffentlicht: (2024)
von: Kundu, Niloy Kumar, et al.
Veröffentlicht: (2024)
A High-Fidelity Speech Super Resolution Network using a Complex Global Attention Module with Spectro-Temporal Loss
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025)
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025)
Fake Speech Wild: Detecting Deepfake Speech on Social Media Platform
von: Xie, Yuankun, et al.
Veröffentlicht: (2025)
von: Xie, Yuankun, et al.
Veröffentlicht: (2025)
Improvement and Implementation of a Speech Emotion Recognition Model Based on Dual-Layer LSTM
von: Yang, Xiaoran, et al.
Veröffentlicht: (2024)
von: Yang, Xiaoran, et al.
Veröffentlicht: (2024)
PTS-SNN: A Prompt-Tuned Temporal Shift Spiking Neural Networks for Efficient Speech Emotion Recognition
von: Su, Xun, et al.
Veröffentlicht: (2026)
von: Su, Xun, et al.
Veröffentlicht: (2026)
Color-based Emotion Representation for Speech Emotion Recognition
von: Nagase, Ryotaro, et al.
Veröffentlicht: (2026)
von: Nagase, Ryotaro, et al.
Veröffentlicht: (2026)
Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models
von: Rakotoarivony, Lucas
Veröffentlicht: (2026)
von: Rakotoarivony, Lucas
Veröffentlicht: (2026)
SpeechQualityLLM: LLM-Based Multimodal Assessment of Speech Quality
von: Monjur, Mahathir, et al.
Veröffentlicht: (2025)
von: Monjur, Mahathir, et al.
Veröffentlicht: (2025)
Semi-Supervised Diseased Detection from Speech Dialogues with Multi-Level Data Modeling
von: Li, Xingyuan, et al.
Veröffentlicht: (2026)
von: Li, Xingyuan, et al.
Veröffentlicht: (2026)
GSRM: Generative Speech Reward Model for Speech RLHF
von: Shen, Maohao, et al.
Veröffentlicht: (2026)
von: Shen, Maohao, et al.
Veröffentlicht: (2026)
A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
von: Wang, Jingyuan, et al.
Veröffentlicht: (2024)
von: Wang, Jingyuan, et al.
Veröffentlicht: (2024)
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2026)
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SS-DPPN: A self-supervised dual-path foundation model for the generalizable cardiac audio representation
von: Muna, Ummy Maria, et al.
Veröffentlicht: (2025) -
Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla
von: Ridoy, Md Sazzadul Islam, et al.
Veröffentlicht: (2025) -
Zero-Shot to Zero-Lies: Detecting Bengali Deepfake Audio through Transfer Learning
von: Samu, Most. Sharmin Sultana, et al.
Veröffentlicht: (2025) -
An Investigation Into Various Approaches For Bengali Long-Form Speech Transcription and Bengali Speaker Diarization
von: Jahan, Epshita, et al.
Veröffentlicht: (2026) -
Explainable Multi-Modal Deep Learning for Automatic Detection of Lung Diseases from Respiratory Audio Signals
von: Saky, S M Asiful Islam, et al.
Veröffentlicht: (2025)