Infant Cry Emotion Recognition Using Improved ECAPA-TDNN with Multiscale Feature Fusion and Attention Enhancement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Junyu, Li, Yanxiong, Yu, Haolin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
von: Kunešová, Marie, et al.
Veröffentlicht: (2025)
von: Kunešová, Marie, et al.
Veröffentlicht: (2025)
MGFF-TDNN: A Multi-Granularity Feature Fusion TDNN Model with Depth-Wise Separable Module for Speaker Verification
von: Li, Ya, et al.
Veröffentlicht: (2025)
von: Li, Ya, et al.
Veröffentlicht: (2025)
Enhancing Infant Crying Detection with Gradient Boosting for Improved Emotional and Mental Health Diagnostics
von: Lee, Kyunghun, et al.
Veröffentlicht: (2024)
von: Lee, Kyunghun, et al.
Veröffentlicht: (2024)
Layer-aware TDNN: Speaker Recognition Using Multi-Layer Features from Pre-Trained Models
von: Kim, Jin Sob, et al.
Veröffentlicht: (2024)
von: Kim, Jin Sob, et al.
Veröffentlicht: (2024)
ECAPA2: A Hybrid Neural Network Architecture and Training Strategy for Robust Speaker Embeddings
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2024)
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2024)
Effective Modeling of Critical Contextual Information for TDNN-based Speaker Verification
von: Weng, Shilong, et al.
Veröffentlicht: (2025)
von: Weng, Shilong, et al.
Veröffentlicht: (2025)
ICSD: An Open-source Dataset for Infant Cry and Snoring Detection
von: Liu, Qingyu, et al.
Veröffentlicht: (2024)
von: Liu, Qingyu, et al.
Veröffentlicht: (2024)
Low-Complexity Acoustic Scene Classification Using Parallel Attention-Convolution Network
von: Li, Yanxiong, et al.
Veröffentlicht: (2024)
von: Li, Yanxiong, et al.
Veröffentlicht: (2024)
Leveraging Cross-Attention Transformer and Multi-Feature Fusion for Cross-Linguistic Speech Emotion Recognition
von: Zhao, Ruoyu, et al.
Veröffentlicht: (2025)
von: Zhao, Ruoyu, et al.
Veröffentlicht: (2025)
How Attention Shapes Emotion: A Comparative Study of Attention Mechanisms for Speech Emotion Recognition
von: Casals-Salvador, Marc, et al.
Veröffentlicht: (2026)
von: Casals-Salvador, Marc, et al.
Veröffentlicht: (2026)
LPGNet: A Lightweight Network with Parallel Attention and Gated Fusion for Multimodal Emotion Recognition
von: He, Zhining, et al.
Veröffentlicht: (2025)
von: He, Zhining, et al.
Veröffentlicht: (2025)
LMU-Based Sequential Learning and Posterior Ensemble Fusion for Cross-Domain Infant Cry Classification
von: Jazaeri, Niloofar, et al.
Veröffentlicht: (2026)
von: Jazaeri, Niloofar, et al.
Veröffentlicht: (2026)
NeXt-TDNN: Modernizing Multi-Scale Temporal Convolution Backbone for Speaker Verification
von: Heo, Hyun-Jun, et al.
Veröffentlicht: (2023)
von: Heo, Hyun-Jun, et al.
Veröffentlicht: (2023)
Speech Emotion Recognition Via CNN-Transformer and Multidimensional Attention Mechanism
von: Tang, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Tang, Xiaoyu, et al.
Veröffentlicht: (2024)
Fully Few-shot Class-incremental Audio Classification Using Multi-level Embedding Extractor and Ridge Regression Classifier
von: Si, Yongjie, et al.
Veröffentlicht: (2025)
von: Si, Yongjie, et al.
Veröffentlicht: (2025)
Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model
von: Ueda, Lucas, et al.
Veröffentlicht: (2025)
von: Ueda, Lucas, et al.
Veröffentlicht: (2025)
Improving Audio Question Answering with Variational Inference
von: Chen, Haolin
Veröffentlicht: (2026)
von: Chen, Haolin
Veröffentlicht: (2026)
MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
BSS-CFFMA: Cross-Domain Feature Fusion and Multi-Attention Speech Enhancement Network based on Self-Supervised Embedding
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2024)
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2024)
Fully Few-shot Class-incremental Audio Classification Using Expandable Dual-embedding Extractor
von: Si, Yongjie, et al.
Veröffentlicht: (2024)
von: Si, Yongjie, et al.
Veröffentlicht: (2024)
BSC-UPC at EmoSPeech-IberLEF2024: Attention Pooling for Emotion Recognition
von: Casals-Salvador, Marc, et al.
Veröffentlicht: (2024)
von: Casals-Salvador, Marc, et al.
Veröffentlicht: (2024)
CryCeleb: A Speaker Verification Dataset Based on Infant Cry Sounds
von: Budaghyan, David, et al.
Veröffentlicht: (2023)
von: Budaghyan, David, et al.
Veröffentlicht: (2023)
Speech Enhancement with Overlapped-Frame Information Fusion and Causal Self-Attention
von: Zhang, Yuewei, et al.
Veröffentlicht: (2025)
von: Zhang, Yuewei, et al.
Veröffentlicht: (2025)
Emotion Recognition in Multi-Speaker Conversations through Speaker Identification, Knowledge Distillation, and Hierarchical Fusion
von: Li, Xiao, et al.
Veröffentlicht: (2025)
von: Li, Xiao, et al.
Veröffentlicht: (2025)
Speech-Based Estimation of Schizophrenia Severity Using Feature Fusion
von: Premananth, Gowtham, et al.
Veröffentlicht: (2024)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2024)
Feature Selection via Graph Topology Inference for Soundscape Emotion Recognition
von: Rey, Samuel, et al.
Veröffentlicht: (2025)
von: Rey, Samuel, et al.
Veröffentlicht: (2025)
Using Songs to Improve Kazakh Automatic Speech Recognition
von: Yeshpanov, Rustem
Veröffentlicht: (2026)
von: Yeshpanov, Rustem
Veröffentlicht: (2026)
Heterogeneous Space Fusion and Dual-Dimension Attention: A New Paradigm for Speech Enhancement
von: Zheng, Tao, et al.
Veröffentlicht: (2024)
von: Zheng, Tao, et al.
Veröffentlicht: (2024)
Multi-Scale Temporal Transformer For Speech Emotion Recognition
von: Li, Zhipeng, et al.
Veröffentlicht: (2024)
von: Li, Zhipeng, et al.
Veröffentlicht: (2024)
Mel-FullSubNet: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
Tracking Listener Attention: Gaze-Guided Audio-Visual Speech Enhancement Framework
von: Yang, Hsiang-Cheng, et al.
Veröffentlicht: (2026)
von: Yang, Hsiang-Cheng, et al.
Veröffentlicht: (2026)
Speech Emotion Recognition with ASR Integration
von: Li, Yuanchao
Veröffentlicht: (2026)
von: Li, Yuanchao
Veröffentlicht: (2026)
Dense-TSNet: Dense Connected Two-Stage Structure for Ultra-Lightweight Speech Enhancement
von: Lin, Zizhen, et al.
Veröffentlicht: (2024)
von: Lin, Zizhen, et al.
Veröffentlicht: (2024)
Magnitude and Phase-based Feature Fusion Using Co-attention Mechanism for Speaker recognition
von: Su, Rongfeng, et al.
Veröffentlicht: (2025)
von: Su, Rongfeng, et al.
Veröffentlicht: (2025)
Double Multi-Head Attention Multimodal System for Odyssey 2024 Speech Emotion Recognition Challenge
von: Costa, Federico, et al.
Veröffentlicht: (2024)
von: Costa, Federico, et al.
Veröffentlicht: (2024)
LightCAM: A Fast and Light Implementation of Context-Aware Masking based D-TDNN for Speaker Verification
von: Cao, Di, et al.
Veröffentlicht: (2024)
von: Cao, Di, et al.
Veröffentlicht: (2024)
A Low-rank Matching Attention based Cross-modal Feature Fusion Method for Conversational Emotion Recognition
von: Shou, Yuntao, et al.
Veröffentlicht: (2023)
von: Shou, Yuntao, et al.
Veröffentlicht: (2023)
Bimodal Connection Attention Fusion for Speech Emotion Recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
Leveraging Content and Acoustic Representations for Speech Emotion Recognition
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
von: Kunešová, Marie, et al.
Veröffentlicht: (2025) -
MGFF-TDNN: A Multi-Granularity Feature Fusion TDNN Model with Depth-Wise Separable Module for Speaker Verification
von: Li, Ya, et al.
Veröffentlicht: (2025) -
Enhancing Infant Crying Detection with Gradient Boosting for Improved Emotional and Mental Health Diagnostics
von: Lee, Kyunghun, et al.
Veröffentlicht: (2024) -
Layer-aware TDNN: Speaker Recognition Using Multi-Layer Features from Pre-Trained Models
von: Kim, Jin Sob, et al.
Veröffentlicht: (2024) -
ECAPA2: A Hybrid Neural Network Architecture and Training Strategy for Robust Speaker Embeddings
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2024)