On combining acoustic and modulation spectrograms in an attention LSTM-based system for speech intelligibility level classification
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Gallardo-Antolín, Ascensión, Montero, Juan M. |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
An Attention Long Short-Term Memory based system for automatic classification of speech intelligibility
par: Fernández-Díaz, Miguel, et autres
Publié: (2024)
par: Fernández-Díaz, Miguel, et autres
Publié: (2024)
Enhancement of a Text-Independent Speaker Verification System by using Feature Combination and Parallel-Structure Classifiers
par: Abdalmalak, Kerlos Atia, et autres
Publié: (2024)
par: Abdalmalak, Kerlos Atia, et autres
Publié: (2024)
Multitaper mel-spectrograms for keyword spotting
par: de Souza, Douglas Baptista, et autres
Publié: (2024)
par: de Souza, Douglas Baptista, et autres
Publié: (2024)
Automatic Detection of Depression in Speech Using Ensemble Convolutional Neural Networks
par: Vázquez-Romero, Adrián, et autres
Publié: (2024)
par: Vázquez-Romero, Adrián, et autres
Publié: (2024)
Throat and acoustic paired speech dataset for deep learning-based speech enhancement
par: Kim, Yunsik, et autres
Publié: (2025)
par: Kim, Yunsik, et autres
Publié: (2025)
Comparison of spectrogram scaling in multi-label Music Genre Recognition
par: Karpiński, Bartosz, et autres
Publié: (2025)
par: Karpiński, Bartosz, et autres
Publié: (2025)
Fusion approaches for emotion recognition from speech using acoustic and text-based features
par: Pepino, Leonardo, et autres
Publié: (2024)
par: Pepino, Leonardo, et autres
Publié: (2024)
Dementia classification from spontaneous speech using wrapper-based feature selection
par: Niemelä, Marko, et autres
Publié: (2025)
par: Niemelä, Marko, et autres
Publié: (2025)
A low latency attention module for streaming self-supervised speech representation learning
par: Ma, Jianbo, et autres
Publié: (2023)
par: Ma, Jianbo, et autres
Publié: (2023)
On the social bias of speech self-supervised models
par: Lin, Yi-Cheng, et autres
Publié: (2024)
par: Lin, Yi-Cheng, et autres
Publié: (2024)
MelHuBERT: A simplified HuBERT on Mel spectrograms
par: Lin, Tzu-Quan, et autres
Publié: (2022)
par: Lin, Tzu-Quan, et autres
Publié: (2022)
Repurposing Image Diffusion Models for Training-Free Music Style Transfer on Mel-spectrograms
par: Wang, Heehwan, et autres
Publié: (2024)
par: Wang, Heehwan, et autres
Publié: (2024)
Exploring speech style spaces with language models: Emotional TTS without emotion labels
par: Chandra, Shreeram Suresh, et autres
Publié: (2024)
par: Chandra, Shreeram Suresh, et autres
Publié: (2024)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
par: Maiti, Soumi, et autres
Publié: (2023)
par: Maiti, Soumi, et autres
Publié: (2023)
Selfsupervised learning for pathological speech detection
par: Sheikh, Shakeel Ahmad
Publié: (2024)
par: Sheikh, Shakeel Ahmad
Publié: (2024)
Towards the Synthesis of Non-speech Vocalizations
par: Hoq, Enjamamul, et autres
Publié: (2024)
par: Hoq, Enjamamul, et autres
Publié: (2024)
Treble10: A high-quality dataset for far-field speech recognition, dereverberation, and enhancement
par: Mullins, Sarabeth S., et autres
Publié: (2025)
par: Mullins, Sarabeth S., et autres
Publié: (2025)
Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
par: Bataev, Vladimir, et autres
Publié: (2023)
par: Bataev, Vladimir, et autres
Publié: (2023)
Room-acoustic simulations as an alternative to measurements for audio-algorithm evaluation
par: Götz, Georg, et autres
Publié: (2025)
par: Götz, Georg, et autres
Publié: (2025)
Modeling speech emotion with label variance and analyzing performance across speakers and unseen acoustic conditions
par: Mitra, Vikramjit, et autres
Publié: (2025)
par: Mitra, Vikramjit, et autres
Publié: (2025)
Towards objective and interpretable speech disorder assessment: a comparative analysis of CNN and transformer-based models
par: Maisonneuve, Malo, et autres
Publié: (2024)
par: Maisonneuve, Malo, et autres
Publié: (2024)
MPIPN: A Multi Physics-Informed PointNet for solving parametric acoustic-structure systems
par: Wang, Chu, et autres
Publié: (2024)
par: Wang, Chu, et autres
Publié: (2024)
The evaluation of a code-switched Sepedi-English automatic speech recognition system
par: Phaladi, Amanda, et autres
Publié: (2024)
par: Phaladi, Amanda, et autres
Publié: (2024)
Incremental learning for audio classification with Hebbian Deep Neural Networks
par: Casciotti, Riccardo, et autres
Publié: (2026)
par: Casciotti, Riccardo, et autres
Publié: (2026)
CR-CTC: Consistency regularization on CTC for improved speech recognition
par: Yao, Zengwei, et autres
Publié: (2024)
par: Yao, Zengwei, et autres
Publié: (2024)
Single-channel speech enhancement using learnable loss mixup
par: Chang, Oscar, et autres
Publié: (2023)
par: Chang, Oscar, et autres
Publié: (2023)
Zipformer: A faster and better encoder for automatic speech recognition
par: Yao, Zengwei, et autres
Publié: (2023)
par: Yao, Zengwei, et autres
Publié: (2023)
Robustifying automatic speech recognition by extracting slowly varying features
par: Pizarro, Matías, et autres
Publié: (2021)
par: Pizarro, Matías, et autres
Publié: (2021)
Omni-directional attention mechanism based on Mamba for speech separation
par: Xue, Ke, et autres
Publié: (2026)
par: Xue, Ke, et autres
Publié: (2026)
Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture
par: Ouyang, Qianhe
Publié: (2025)
par: Ouyang, Qianhe
Publié: (2025)
A contrastive-learning approach for auditory attention detection
par: Bajestan, Seyed Ali Alavi, et autres
Publié: (2024)
par: Bajestan, Seyed Ali Alavi, et autres
Publié: (2024)
Deep learning classification system for coconut maturity levels based on acoustic signals
par: Caladcad, June Anne, et autres
Publié: (2024)
par: Caladcad, June Anne, et autres
Publié: (2024)
U-Mamba-Net: A highly efficient Mamba-based U-net style network for noisy and reverberant speech separation
par: Dang, Shaoxiang, et autres
Publié: (2024)
par: Dang, Shaoxiang, et autres
Publié: (2024)
Late fusion ensembles for speech recognition on diverse input audio representations
par: Jezidžić, Marin, et autres
Publié: (2024)
par: Jezidžić, Marin, et autres
Publié: (2024)
Boosting keyword spotting through on-device learnable user speech characteristics
par: Cioflan, Cristian, et autres
Publié: (2024)
par: Cioflan, Cristian, et autres
Publié: (2024)
Generalizable speech deepfake detection via meta-learned LoRA
par: Laakkonen, Janne, et autres
Publié: (2025)
par: Laakkonen, Janne, et autres
Publié: (2025)
An LSTM-Based Chord Generation System Using Chroma Histogram Representations
par: Hardwick, Jack
Publié: (2024)
par: Hardwick, Jack
Publié: (2024)
Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
par: Leygue, Tahitoa, et autres
Publié: (2025)
par: Leygue, Tahitoa, et autres
Publié: (2025)
Acoustic characterization of speech rhythm: going beyond metrics with recurrent neural networks
par: Deloche, François, et autres
Publié: (2024)
par: Deloche, François, et autres
Publié: (2024)
Context-aware child-directed speech detection from long-form recordings
par: Charlot, Théo, et autres
Publié: (2026)
par: Charlot, Théo, et autres
Publié: (2026)
Documents similaires
-
An Attention Long Short-Term Memory based system for automatic classification of speech intelligibility
par: Fernández-Díaz, Miguel, et autres
Publié: (2024) -
Enhancement of a Text-Independent Speaker Verification System by using Feature Combination and Parallel-Structure Classifiers
par: Abdalmalak, Kerlos Atia, et autres
Publié: (2024) -
Multitaper mel-spectrograms for keyword spotting
par: de Souza, Douglas Baptista, et autres
Publié: (2024) -
Automatic Detection of Depression in Speech Using Ensemble Convolutional Neural Networks
par: Vázquez-Romero, Adrián, et autres
Publié: (2024) -
Throat and acoustic paired speech dataset for deep learning-based speech enhancement
par: Kim, Yunsik, et autres
Publié: (2025)