Throat and acoustic paired speech dataset for deep learning-based speech enhancement
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Yunsik, Song, Yonghun, Chung, Yoonyoung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Selfsupervised learning for pathological speech detection
di: Sheikh, Shakeel Ahmad
Pubblicazione: (2024)
di: Sheikh, Shakeel Ahmad
Pubblicazione: (2024)
Fusion approaches for emotion recognition from speech using acoustic and text-based features
di: Pepino, Leonardo, et al.
Pubblicazione: (2024)
di: Pepino, Leonardo, et al.
Pubblicazione: (2024)
Single-channel speech enhancement using learnable loss mixup
di: Chang, Oscar, et al.
Pubblicazione: (2023)
di: Chang, Oscar, et al.
Pubblicazione: (2023)
Generalizable speech deepfake detection via meta-learned LoRA
di: Laakkonen, Janne, et al.
Pubblicazione: (2025)
di: Laakkonen, Janne, et al.
Pubblicazione: (2025)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
di: Maiti, Soumi, et al.
Pubblicazione: (2023)
di: Maiti, Soumi, et al.
Pubblicazione: (2023)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
Towards the Synthesis of Non-speech Vocalizations
di: Hoq, Enjamamul, et al.
Pubblicazione: (2024)
di: Hoq, Enjamamul, et al.
Pubblicazione: (2024)
Unsupervised speech enhancement with spectral kurtosis and double deep priors
di: Ohnaka, Hien, et al.
Pubblicazione: (2024)
di: Ohnaka, Hien, et al.
Pubblicazione: (2024)
Monaural speech enhancement on drone via Adapter based transfer learning
di: Chen, Xingyu, et al.
Pubblicazione: (2024)
di: Chen, Xingyu, et al.
Pubblicazione: (2024)
Objective and subjective evaluation of speech enhancement methods in the UDASE task of the 7th CHiME challenge
di: Leglaive, Simon, et al.
Pubblicazione: (2024)
di: Leglaive, Simon, et al.
Pubblicazione: (2024)
Dementia classification from spontaneous speech using wrapper-based feature selection
di: Niemelä, Marko, et al.
Pubblicazione: (2025)
di: Niemelä, Marko, et al.
Pubblicazione: (2025)
Real-time multichannel deep speech enhancement in hearing aids: Comparing monaural and binaural processing in complex acoustic scenarios
di: Westhausen, Nils L., et al.
Pubblicazione: (2024)
di: Westhausen, Nils L., et al.
Pubblicazione: (2024)
A multimodal dynamical variational autoencoder for audiovisual speech representation learning
di: Sadok, Samir, et al.
Pubblicazione: (2023)
di: Sadok, Samir, et al.
Pubblicazione: (2023)
An Attention Long Short-Term Memory based system for automatic classification of speech intelligibility
di: Fernández-Díaz, Miguel, et al.
Pubblicazione: (2024)
di: Fernández-Díaz, Miguel, et al.
Pubblicazione: (2024)
Zipformer: A faster and better encoder for automatic speech recognition
di: Yao, Zengwei, et al.
Pubblicazione: (2023)
di: Yao, Zengwei, et al.
Pubblicazione: (2023)
CR-CTC: Consistency regularization on CTC for improved speech recognition
di: Yao, Zengwei, et al.
Pubblicazione: (2024)
di: Yao, Zengwei, et al.
Pubblicazione: (2024)
Robustifying automatic speech recognition by extracting slowly varying features
di: Pizarro, Matías, et al.
Pubblicazione: (2021)
di: Pizarro, Matías, et al.
Pubblicazione: (2021)
Late fusion ensembles for speech recognition on diverse input audio representations
di: Jezidžić, Marin, et al.
Pubblicazione: (2024)
di: Jezidžić, Marin, et al.
Pubblicazione: (2024)
Boosting keyword spotting through on-device learnable user speech characteristics
di: Cioflan, Cristian, et al.
Pubblicazione: (2024)
di: Cioflan, Cristian, et al.
Pubblicazione: (2024)
Towards objective and interpretable speech disorder assessment: a comparative analysis of CNN and transformer-based models
di: Maisonneuve, Malo, et al.
Pubblicazione: (2024)
di: Maisonneuve, Malo, et al.
Pubblicazione: (2024)
Introduction to speech recognition
di: Dauphin, Gabriel
Pubblicazione: (2024)
di: Dauphin, Gabriel
Pubblicazione: (2024)
An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech
di: Deng, Qingkun, et al.
Pubblicazione: (2024)
di: Deng, Qingkun, et al.
Pubblicazione: (2024)
Context-aware child-directed speech detection from long-form recordings
di: Charlot, Théo, et al.
Pubblicazione: (2026)
di: Charlot, Théo, et al.
Pubblicazione: (2026)
Acoustic characterization of speech rhythm: going beyond metrics with recurrent neural networks
di: Deloche, François, et al.
Pubblicazione: (2024)
di: Deloche, François, et al.
Pubblicazione: (2024)
Inter-channel Conv-TasNet for multichannel speech enhancement
di: Lee, Dongheon, et al.
Pubblicazione: (2021)
di: Lee, Dongheon, et al.
Pubblicazione: (2021)
Self-supervised learning of speech representations with Dutch archival data
di: Vaessen, Nik, et al.
Pubblicazione: (2025)
di: Vaessen, Nik, et al.
Pubblicazione: (2025)
Modeling speech emotion with label variance and analyzing performance across speakers and unseen acoustic conditions
di: Mitra, Vikramjit, et al.
Pubblicazione: (2025)
di: Mitra, Vikramjit, et al.
Pubblicazione: (2025)
SeMaScore : a new evaluation metric for automatic speech recognition tasks
di: Sasindran, Zitha, et al.
Pubblicazione: (2024)
di: Sasindran, Zitha, et al.
Pubblicazione: (2024)
Acoustic-to-articulatory inversion for dysarthric speech: Are pre-trained self-supervised representations favorable?
di: Maharana, Sarthak Kumar, et al.
Pubblicazione: (2023)
di: Maharana, Sarthak Kumar, et al.
Pubblicazione: (2023)
Online speaker diarization of meetings guided by speech separation
di: Gruttadauria, Elio, et al.
Pubblicazione: (2024)
di: Gruttadauria, Elio, et al.
Pubblicazione: (2024)
U-Mamba-Net: A highly efficient Mamba-based U-net style network for noisy and reverberant speech separation
di: Dang, Shaoxiang, et al.
Pubblicazione: (2024)
di: Dang, Shaoxiang, et al.
Pubblicazione: (2024)
CognoSpeak: an automatic, remote assessment of early cognitive decline in real-world conversational speech
di: Pahar, Madhurananda, et al.
Pubblicazione: (2025)
di: Pahar, Madhurananda, et al.
Pubblicazione: (2025)
GDiffuSE: Diffusion-based speech enhancement with noise model guidance
di: Yanir, Efrayim, et al.
Pubblicazione: (2025)
di: Yanir, Efrayim, et al.
Pubblicazione: (2025)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2025)
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2025)
A vector quantized masked autoencoder for audiovisual speech emotion recognition
di: Sadok, Samir, et al.
Pubblicazione: (2023)
di: Sadok, Samir, et al.
Pubblicazione: (2023)
A low latency attention module for streaming self-supervised speech representation learning
di: Ma, Jianbo, et al.
Pubblicazione: (2023)
di: Ma, Jianbo, et al.
Pubblicazione: (2023)
Towards measuring fairness in speech recognition: Fair-Speech dataset
di: Veliche, Irina-Elena, et al.
Pubblicazione: (2024)
di: Veliche, Irina-Elena, et al.
Pubblicazione: (2024)
SPGM: Prioritizing Local Features for enhanced speech separation performance
di: Yip, Jia Qi, et al.
Pubblicazione: (2023)
di: Yip, Jia Qi, et al.
Pubblicazione: (2023)
Detecting Throat Cancer from Speech Signals using Machine Learning: A Scoping Literature Review
di: Paterson, Mary, et al.
Pubblicazione: (2023)
di: Paterson, Mary, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Selfsupervised learning for pathological speech detection
di: Sheikh, Shakeel Ahmad
Pubblicazione: (2024) -
Fusion approaches for emotion recognition from speech using acoustic and text-based features
di: Pepino, Leonardo, et al.
Pubblicazione: (2024) -
Single-channel speech enhancement using learnable loss mixup
di: Chang, Oscar, et al.
Pubblicazione: (2023) -
Generalizable speech deepfake detection via meta-learned LoRA
di: Laakkonen, Janne, et al.
Pubblicazione: (2025) -
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
di: Maiti, Soumi, et al.
Pubblicazione: (2023)