EmoAugNet: A Signal-Augmented Hybrid CNN-LSTM Framework for Speech Emotion Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Paul, Durjoy Chandra, Saha, Gaurob, Hossain, Md Amjad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Human Feedback Driven Dynamic Speech Emotion Recognition
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025)
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025)
Artificial Neural Networks to Recognize Speakers Division from Continuous Bengali Speech
von: Ali, Hasmot, et al.
Veröffentlicht: (2024)
von: Ali, Hasmot, et al.
Veröffentlicht: (2024)
STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition
von: Chang, Yi, et al.
Veröffentlicht: (2024)
von: Chang, Yi, et al.
Veröffentlicht: (2024)
Emotion-Disentangled Embedding Alignment for Noise-Robust and Cross-Corpus Speech Emotion Recognition
von: Tiwari, Upasana, et al.
Veröffentlicht: (2025)
von: Tiwari, Upasana, et al.
Veröffentlicht: (2025)
Optimizing Multilingual Text-To-Speech with Accents & Emotions
von: Pawar, Pranav, et al.
Veröffentlicht: (2025)
von: Pawar, Pranav, et al.
Veröffentlicht: (2025)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers
von: Mishra, Ruchik, et al.
Veröffentlicht: (2024)
von: Mishra, Ruchik, et al.
Veröffentlicht: (2024)
LSTM-CNN Network for Audio Signature Analysis in Noisy Environments
von: Damacharla, Praveen, et al.
Veröffentlicht: (2023)
von: Damacharla, Praveen, et al.
Veröffentlicht: (2023)
Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection
von: Chandra, Joydeep
Veröffentlicht: (2026)
von: Chandra, Joydeep
Veröffentlicht: (2026)
Early Detection of Furniture-Infesting Wood-Boring Beetles Using CNN-LSTM Networks and MFCC-Based Acoustic Features
von: Manukalpa, J. M. Chan Sri, et al.
Veröffentlicht: (2025)
von: Manukalpa, J. M. Chan Sri, et al.
Veröffentlicht: (2025)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
EmoHeal: An End-to-End System for Personalized Therapeutic Music Retrieval from Fine-grained Emotions
von: Wan, Xinchen, et al.
Veröffentlicht: (2025)
von: Wan, Xinchen, et al.
Veröffentlicht: (2025)
Arabic Little STT: Arabic Children Speech Recognition Dataset
von: Alkadri, Mouhand, et al.
Veröffentlicht: (2025)
von: Alkadri, Mouhand, et al.
Veröffentlicht: (2025)
Morse Code-Enabled Speech Recognition for Individuals with Visual and Hearing Impairments
von: Choudhury, Ritabrata Roy
Veröffentlicht: (2024)
von: Choudhury, Ritabrata Roy
Veröffentlicht: (2024)
EmoKnob: Enhance Voice Cloning with Fine-Grained Emotion Control
von: Chen, Haozhe, et al.
Veröffentlicht: (2024)
von: Chen, Haozhe, et al.
Veröffentlicht: (2024)
Three-Class Emotion Classification for Audiovisual Scenes Based on Ensemble Learning Scheme
von: Xiong, Xiangrui, et al.
Veröffentlicht: (2025)
von: Xiong, Xiangrui, et al.
Veröffentlicht: (2025)
Improved Dysarthric Speech to Text Conversion via TTS Personalization
von: Mihajlik, Péter, et al.
Veröffentlicht: (2025)
von: Mihajlik, Péter, et al.
Veröffentlicht: (2025)
LightBeam: An Accurate and Memory-Efficient CTC Decoder for Speech Neuroprostheses
von: Feghhi, Ebrahim, et al.
Veröffentlicht: (2026)
von: Feghhi, Ebrahim, et al.
Veröffentlicht: (2026)
Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
von: Zhao, Zhixian, et al.
Veröffentlicht: (2024)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2024)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization
von: Dementyev, Artem, et al.
Veröffentlicht: (2025)
von: Dementyev, Artem, et al.
Veröffentlicht: (2025)
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
von: Sankey-Olsen, Cuno, et al.
Veröffentlicht: (2025)
von: Sankey-Olsen, Cuno, et al.
Veröffentlicht: (2025)
Adaptation and Optimization of Automatic Speech Recognition (ASR) for the Maritime Domain in the Field of VHF Communication
von: Nakilcioglu, Emin Cagatay, et al.
Veröffentlicht: (2023)
von: Nakilcioglu, Emin Cagatay, et al.
Veröffentlicht: (2023)
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
von: Nowrin, Sadia, et al.
Veröffentlicht: (2024)
von: Nowrin, Sadia, et al.
Veröffentlicht: (2024)
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
von: Dutta, Satwik, et al.
Veröffentlicht: (2025)
von: Dutta, Satwik, et al.
Veröffentlicht: (2025)
Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
von: Han, Zhichen, et al.
Veröffentlicht: (2024)
von: Han, Zhichen, et al.
Veröffentlicht: (2024)
Enhanced Speech Emotion Recognition with Efficient Channel Attention Guided Deep CNN-BiLSTM Framework
von: Kundu, Niloy Kumar, et al.
Veröffentlicht: (2024)
von: Kundu, Niloy Kumar, et al.
Veröffentlicht: (2024)
Live Music Models
von: Lyria Team, et al.
Veröffentlicht: (2025)
von: Lyria Team, et al.
Veröffentlicht: (2025)
Apollo: An Interactive Environment for Generating Symbolic Musical Phrases using Corpus-based Style Imitation
von: Tchemeube, Renaud Bougueng, et al.
Veröffentlicht: (2025)
von: Tchemeube, Renaud Bougueng, et al.
Veröffentlicht: (2025)
Calliope: An Online Generative Music System for Symbolic Multi-Track Composition
von: Tchemeube, Renaud Bougueng, et al.
Veröffentlicht: (2025)
von: Tchemeube, Renaud Bougueng, et al.
Veröffentlicht: (2025)
ArabEmoNet: A Lightweight Hybrid 2D CNN-BiLSTM Model with Attention for Robust Arabic Speech Emotion Recognition
von: Abouzeid, Ali, et al.
Veröffentlicht: (2025)
von: Abouzeid, Ali, et al.
Veröffentlicht: (2025)
Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models
von: Dietrich, Juergen
Veröffentlicht: (2026)
von: Dietrich, Juergen
Veröffentlicht: (2026)
VoiceX: A Text-To-Speech Framework for Custom Voices
von: Mertes, Silvan, et al.
Veröffentlicht: (2024)
von: Mertes, Silvan, et al.
Veröffentlicht: (2024)
SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
von: Yuan, Kuang, et al.
Veröffentlicht: (2025)
von: Yuan, Kuang, et al.
Veröffentlicht: (2025)
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
von: Shi, Weiyan, et al.
Veröffentlicht: (2025)
von: Shi, Weiyan, et al.
Veröffentlicht: (2025)
Sound-Based Recognition of Touch Gestures and Emotions for Enhanced Human-Robot Interaction
von: Hou, Yuanbo, et al.
Veröffentlicht: (2024)
von: Hou, Yuanbo, et al.
Veröffentlicht: (2024)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
von: Nishida, Naoto, et al.
Veröffentlicht: (2025)
von: Nishida, Naoto, et al.
Veröffentlicht: (2025)
A Multimodal Emotion Recognition System: Integrating Facial Expressions, Body Movement, Speech, and Spoken Language
von: Kraack, Kris
Veröffentlicht: (2024)
von: Kraack, Kris
Veröffentlicht: (2024)
Enhancing Generalization in PPG-Based Emotion Measurement with a CNN-TCN-LSTM Model
von: Alghoul, Karim, et al.
Veröffentlicht: (2025)
von: Alghoul, Karim, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Human Feedback Driven Dynamic Speech Emotion Recognition
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025) -
Artificial Neural Networks to Recognize Speakers Division from Continuous Bengali Speech
von: Ali, Hasmot, et al.
Veröffentlicht: (2024) -
STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition
von: Chang, Yi, et al.
Veröffentlicht: (2024) -
Emotion-Disentangled Embedding Alignment for Noise-Robust and Cross-Corpus Speech Emotion Recognition
von: Tiwari, Upasana, et al.
Veröffentlicht: (2025) -
Optimizing Multilingual Text-To-Speech with Accents & Emotions
von: Pawar, Pranav, et al.
Veröffentlicht: (2025)