Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lau, Hok-Shing, Huntly, Mark, Morgan, Nathon, Iyenoma, Adesua, Zeng, Biao, Bashford, Tim |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
par: Aldeneh, Zakaria, et autres
Publié: (2024)
par: Aldeneh, Zakaria, et autres
Publié: (2024)
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
par: Ku, Pin-Jui, et autres
Publié: (2024)
par: Ku, Pin-Jui, et autres
Publié: (2024)
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
par: Zheng, Xiuwen, et autres
Publié: (2024)
par: Zheng, Xiuwen, et autres
Publié: (2024)
Private kNN-VC: Interpretable Anonymization of Converted Speech
par: Franzreb, Carlos, et autres
Publié: (2025)
par: Franzreb, Carlos, et autres
Publié: (2025)
Active Learning of Non-semantic Speech Tasks with Pretrained Models
par: Lee, Harlin, et autres
Publié: (2022)
par: Lee, Harlin, et autres
Publié: (2022)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
par: Huo, Mingyue, et autres
Publié: (2025)
par: Huo, Mingyue, et autres
Publié: (2025)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
par: Nespoli, Francesco, et autres
Publié: (2024)
par: Nespoli, Francesco, et autres
Publié: (2024)
Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module
par: Cui, Zhongjian, et autres
Publié: (2025)
par: Cui, Zhongjian, et autres
Publié: (2025)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
par: Tian, Jingguang, et autres
Publié: (2024)
par: Tian, Jingguang, et autres
Publié: (2024)
Pitch Accent Detection improves Pretrained Automatic Speech Recognition
par: Sasu, David, et autres
Publié: (2025)
par: Sasu, David, et autres
Publié: (2025)
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
par: Wang, Hui, et autres
Publié: (2025)
par: Wang, Hui, et autres
Publié: (2025)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
par: Farhadipour, Aref, et autres
Publié: (2024)
par: Farhadipour, Aref, et autres
Publié: (2024)
Amplifying Artifacts with Speech Enhancement in Voice Anti-spoofing
par: Trachu, Thanapat, et autres
Publié: (2025)
par: Trachu, Thanapat, et autres
Publié: (2025)
Speech Synthesis along Perceptual Voice Quality Dimensions
par: Rautenberg, Frederik, et autres
Publié: (2025)
par: Rautenberg, Frederik, et autres
Publié: (2025)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
par: Li, Jialu, et autres
Publié: (2024)
par: Li, Jialu, et autres
Publié: (2024)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
par: Alsayegh, Ali, et autres
Publié: (2025)
par: Alsayegh, Ali, et autres
Publié: (2025)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
par: Leung, Wing-Zin, et autres
Publié: (2024)
par: Leung, Wing-Zin, et autres
Publié: (2024)
Quality Assessment of Noisy and Enhanced Speech with Limited Data: UWB-NTIS System for VoiceMOS 2024
par: Kunešová, Marie, et autres
Publié: (2025)
par: Kunešová, Marie, et autres
Publié: (2025)
Self-Supervised Speech Quality Assessment (S3QA): Leveraging Speech Foundation Models for a Scalable Speech Quality Metric
par: Ogg, Mattson, et autres
Publié: (2025)
par: Ogg, Mattson, et autres
Publié: (2025)
Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
par: Byun, Kyungguen, et autres
Publié: (2025)
par: Byun, Kyungguen, et autres
Publié: (2025)
Speech to Speech Synthesis for Voice Impersonation
par: Johnson, Bjorn, et autres
Publié: (2026)
par: Johnson, Bjorn, et autres
Publié: (2026)
Modeling of Speech-dependent Own Voice Transfer Characteristics for Hearables with In-ear Microphones
par: Ohlenbusch, Mattes, et autres
Publié: (2023)
par: Ohlenbusch, Mattes, et autres
Publié: (2023)
Speech-dependent Modeling of Own Voice Transfer Characteristics for In-ear Microphones in Hearables
par: Ohlenbusch, Mattes, et autres
Publié: (2023)
par: Ohlenbusch, Mattes, et autres
Publié: (2023)
On the Relevance of Clinical Assessment Tasks for the Automatic Detection of Parkinson's Disease Medication State from Speech
par: Gimeno-Gómez, David, et autres
Publié: (2025)
par: Gimeno-Gómez, David, et autres
Publié: (2025)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
par: Chen, Peikun, et autres
Publié: (2024)
par: Chen, Peikun, et autres
Publié: (2024)
SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
par: Saeki, Takaaki, et autres
Publié: (2024)
par: Saeki, Takaaki, et autres
Publié: (2024)
Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
par: Zhu, Xiaoxu, et autres
Publié: (2025)
par: Zhu, Xiaoxu, et autres
Publié: (2025)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
par: Li, Xuyuan, et autres
Publié: (2024)
par: Li, Xuyuan, et autres
Publié: (2024)
The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
par: Huang, Wen-Chin, et autres
Publié: (2024)
par: Huang, Wen-Chin, et autres
Publié: (2024)
RAVE for Speech: Efficient Voice Conversion at High Sampling Rates
par: Bargum, Anders R., et autres
Publié: (2024)
par: Bargum, Anders R., et autres
Publié: (2024)
Fine-Grained and Interpretable Neural Speech Editing
par: Morrison, Max, et autres
Publié: (2024)
par: Morrison, Max, et autres
Publié: (2024)
Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
par: Geng, Haopeng, et autres
Publié: (2024)
par: Geng, Haopeng, et autres
Publié: (2024)
Compact Speech Translation Models via Discrete Speech Units Pretraining
par: Lam, Tsz Kin, et autres
Publié: (2024)
par: Lam, Tsz Kin, et autres
Publié: (2024)
Speech Denoising with Auditory Models
par: Saddler, Mark R., et autres
Publié: (2020)
par: Saddler, Mark R., et autres
Publié: (2020)
ZipEnhancer: Dual-Path Down-Up Sampling-based Zipformer for Monaural Speech Enhancement
par: Wang, Haoxu, et autres
Publié: (2025)
par: Wang, Haoxu, et autres
Publié: (2025)
VC-ENHANCE: Speech Restoration with Integrated Noise Suppression and Voice Conversion
par: Byun, Kyungguen, et autres
Publié: (2024)
par: Byun, Kyungguen, et autres
Publié: (2024)
Loudspeaker Beamforming to Enhance Speech Recognition Performance of Voice Driven Applications
par: de Groot, Dimme, et autres
Publié: (2025)
par: de Groot, Dimme, et autres
Publié: (2025)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
par: Park, Nohil, et autres
Publié: (2024)
par: Park, Nohil, et autres
Publié: (2024)
VoiceRestore: Flow-Matching Transformers for Speech Recording Quality Restoration
par: Kirdey, Stanislav
Publié: (2025)
par: Kirdey, Stanislav
Publié: (2025)
Preserving Speaker Information in Direct Speech-to-Speech Translation with Non-Autoregressive Generation and Pretraining
par: Zhou, Rui, et autres
Publié: (2024)
par: Zhou, Rui, et autres
Publié: (2024)
Documents similaires
-
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
par: Aldeneh, Zakaria, et autres
Publié: (2024) -
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
par: Ku, Pin-Jui, et autres
Publié: (2024) -
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
par: Zheng, Xiuwen, et autres
Publié: (2024) -
Private kNN-VC: Interpretable Anonymization of Converted Speech
par: Franzreb, Carlos, et autres
Publié: (2025) -
Active Learning of Non-semantic Speech Tasks with Pretrained Models
par: Lee, Harlin, et autres
Publié: (2022)