Towards objective and interpretable speech disorder assessment: a comparative analysis of CNN and transformer-based models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Maisonneuve, Malo, Fredouille, Corinne, Lalain, Muriel, Ghio, Alain, Woisard, Virginie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech
von: Deng, Qingkun, et al.
Veröffentlicht: (2024)
von: Deng, Qingkun, et al.
Veröffentlicht: (2024)
Towards a generalized monaural and binaural auditory model for psychoacoustics and speech intelligibility
von: Biberger, Thomas, et al.
Veröffentlicht: (2021)
von: Biberger, Thomas, et al.
Veröffentlicht: (2021)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026)
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026)
GDiffuSE: Diffusion-based speech enhancement with noise model guidance
von: Yanir, Efrayim, et al.
Veröffentlicht: (2025)
von: Yanir, Efrayim, et al.
Veröffentlicht: (2025)
Towards robust paralinguistic assessment for real-world mobile health (mHealth) monitoring: an initial study of reverberation effects on speech
von: Dineley, Judith, et al.
Veröffentlicht: (2023)
von: Dineley, Judith, et al.
Veröffentlicht: (2023)
Discrimination loss vs. SRT: A model-based approach towards harmonizing speech test interpretations
von: Buhl, Mareike, et al.
Veröffentlicht: (2025)
von: Buhl, Mareike, et al.
Veröffentlicht: (2025)
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
von: An, Keyu, et al.
Veröffentlicht: (2024)
von: An, Keyu, et al.
Veröffentlicht: (2024)
Language model integration based on memory control for sequence to sequence speech recognition
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
Phoneme-based speech recognition driven by large language models and sampling marginalization
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Learnings from curating a trustworthy, well-annotated, and useful dataset of disordered English speech
von: Jiang, Pan-Pan, et al.
Veröffentlicht: (2024)
von: Jiang, Pan-Pan, et al.
Veröffentlicht: (2024)
Probing mental health information in speech foundation models
von: de Gennes, Marc, et al.
Veröffentlicht: (2024)
von: de Gennes, Marc, et al.
Veröffentlicht: (2024)
WhisperFlow: speech foundation models in real time
von: Wang, Rongxiang, et al.
Veröffentlicht: (2024)
von: Wang, Rongxiang, et al.
Veröffentlicht: (2024)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
von: Pascu, Octavian, et al.
Veröffentlicht: (2023)
von: Pascu, Octavian, et al.
Veröffentlicht: (2023)
Omni-directional attention mechanism based on Mamba for speech separation
von: Xue, Ke, et al.
Veröffentlicht: (2026)
von: Xue, Ke, et al.
Veröffentlicht: (2026)
Adaptive Convolution for CNN-based Speech Enhancement Models
von: Wang, Dahan, et al.
Veröffentlicht: (2025)
von: Wang, Dahan, et al.
Veröffentlicht: (2025)
Graph-based multi-Feature fusion method for speech emotion recognition
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
Monaural speech enhancement on drone via Adapter based transfer learning
von: Chen, Xingyu, et al.
Veröffentlicht: (2024)
von: Chen, Xingyu, et al.
Veröffentlicht: (2024)
Single-channel speech enhancement by using psychoacoustical model inspired fusion framework
von: Samui, Suman
Veröffentlicht: (2022)
von: Samui, Suman
Veröffentlicht: (2022)
Multichannel blind speech source separation with a disjoint constraint source model
von: Wang, Jianyu, et al.
Veröffentlicht: (2024)
von: Wang, Jianyu, et al.
Veröffentlicht: (2024)
Towards interpretable emotion recognition: Identifying key features with machine learning
von: Kaloga, Yacouba, et al.
Veröffentlicht: (2025)
von: Kaloga, Yacouba, et al.
Veröffentlicht: (2025)
On the relationship between speech and hearing
von: Umesh, Srinivasan, et al.
Veröffentlicht: (2024)
von: Umesh, Srinivasan, et al.
Veröffentlicht: (2024)
Building speech corpus with diverse voice characteristics for its prompt-based representation
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
Automatic speech recognition for the Nepali language using CNN, bidirectional LSTM and ResNet
von: Dhakal, Manish, et al.
Veröffentlicht: (2024)
von: Dhakal, Manish, et al.
Veröffentlicht: (2024)
Relationship between objective and subjective perceptual measures of speech in individuals with head and neck cancer
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2025)
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2025)
Determined blind source separation via modeling adjacent frequency band correlations in speech signals
von: Wang, Jianyu, et al.
Veröffentlicht: (2025)
von: Wang, Jianyu, et al.
Veröffentlicht: (2025)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
CNN-based Robust Sound Source Localization with SRP-PHAT for the Extreme Edge
von: Yin, Jun, et al.
Veröffentlicht: (2025)
von: Yin, Jun, et al.
Veröffentlicht: (2025)
An adaptive filter bank based neural network approach for time delay estimation and speech enhancement
von: Ma, Lu
Veröffentlicht: (2025)
von: Ma, Lu
Veröffentlicht: (2025)
Towards the Synthesis of Non-speech Vocalizations
von: Hoq, Enjamamul, et al.
Veröffentlicht: (2024)
von: Hoq, Enjamamul, et al.
Veröffentlicht: (2024)
Enhancing CTC-based speech recognition with diverse modeling units
von: Han, Shiyi, et al.
Veröffentlicht: (2024)
von: Han, Shiyi, et al.
Veröffentlicht: (2024)
A lightweight dual-stage framework for personalized speech enhancement based on DeepFilterNet2
von: Serre, Thomas, et al.
Veröffentlicht: (2024)
von: Serre, Thomas, et al.
Veröffentlicht: (2024)
Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
Distilling a speech and music encoder with task arithmetic
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
Exploring compressibility of transformer based text-to-music (TTM) models
von: Moschopoulos, Vasileios, et al.
Veröffentlicht: (2024)
von: Moschopoulos, Vasileios, et al.
Veröffentlicht: (2024)
Unsupervised speech enhancement with spectral kurtosis and double deep priors
von: Ohnaka, Hien, et al.
Veröffentlicht: (2024)
von: Ohnaka, Hien, et al.
Veröffentlicht: (2024)
BFA: Real-time Multilingual Text-to-speech Forced Alignment
von: Rehman, Abdul, et al.
Veröffentlicht: (2025)
von: Rehman, Abdul, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024) -
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024) -
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025) -
An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech
von: Deng, Qingkun, et al.
Veröffentlicht: (2024) -
Towards a generalized monaural and binaural auditory model for psychoacoustics and speech intelligibility
von: Biberger, Thomas, et al.
Veröffentlicht: (2021)