Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
Fuente:
arXiv
Salvato in:
| Autori principali: | Fu, Szu-Wei, Hung, Kuo-Hsuan, Tsao, Yu, Wang, Yu-Chiang Frank |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Universal Speech Enhancement with Regression and Generative Mamba
di: Chao, Rong, et al.
Pubblicazione: (2025)
di: Chao, Rong, et al.
Pubblicazione: (2025)
An Investigation of Incorporating Mamba for Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2024)
di: Chao, Rong, et al.
Pubblicazione: (2024)
Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
di: Khan, Muhammad Salman, et al.
Pubblicazione: (2024)
di: Khan, Muhammad Salman, et al.
Pubblicazione: (2024)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2025)
di: Chao, Rong, et al.
Pubblicazione: (2025)
CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR
di: Shao, Nian, et al.
Pubblicazione: (2025)
di: Shao, Nian, et al.
Pubblicazione: (2025)
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
di: Ren, Wenze, et al.
Pubblicazione: (2024)
di: Ren, Wenze, et al.
Pubblicazione: (2024)
LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
di: Chen, Chih-Ning, et al.
Pubblicazione: (2026)
di: Chen, Chih-Ning, et al.
Pubblicazione: (2026)
Magnitude-Phase Dual-Path Speech Enhancement Network based on Self-Supervised Embedding and Perceptual Contrast Stretch Boosting
di: Mattursun, Alimjan, et al.
Pubblicazione: (2025)
di: Mattursun, Alimjan, et al.
Pubblicazione: (2025)
The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
di: Huang, Wen-Chin, et al.
Pubblicazione: (2024)
di: Huang, Wen-Chin, et al.
Pubblicazione: (2024)
A Study on Incorporating Whisper for Robust Speech Assessment
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
di: Ren, Wenze, et al.
Pubblicazione: (2024)
di: Ren, Wenze, et al.
Pubblicazione: (2024)
Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models
di: Shah, Neil, et al.
Pubblicazione: (2024)
di: Shah, Neil, et al.
Pubblicazione: (2024)
Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2026)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2026)
QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
di: Wang, Siyin, et al.
Pubblicazione: (2025)
di: Wang, Siyin, et al.
Pubblicazione: (2025)
Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech Edits
di: Huang, Sung-Feng, et al.
Pubblicazione: (2025)
di: Huang, Sung-Feng, et al.
Pubblicazione: (2025)
Self-Supervised Models for Phoneme Recognition: Applications in Children's Speech for Reading Learning
di: Medin, Lucas Block, et al.
Pubblicazione: (2025)
di: Medin, Lucas Block, et al.
Pubblicazione: (2025)
Improving Speech Inversion Through Self-Supervised Embeddings and Enhanced Tract Variables
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2023)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2023)
MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation
di: Chen, Szu-Chi, et al.
Pubblicazione: (2026)
di: Chen, Szu-Chi, et al.
Pubblicazione: (2026)
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
di: Ku, Pin-Jui, et al.
Pubblicazione: (2024)
di: Ku, Pin-Jui, et al.
Pubblicazione: (2024)
Restorative Speech Enhancement: A Progressive Approach Using SE and Codec Modules
di: Chiang, Hsin-Tien, et al.
Pubblicazione: (2024)
di: Chiang, Hsin-Tien, et al.
Pubblicazione: (2024)
DAISY: Data Adaptive Self-Supervised Early Exit for Speech Representation Models
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024)
Improving Voice Quality in Speech Anonymization With Just Perception-Informed Losses
di: Ghosh, Suhita, et al.
Pubblicazione: (2024)
di: Ghosh, Suhita, et al.
Pubblicazione: (2024)
Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
di: Aboeitta, Ahmed, et al.
Pubblicazione: (2025)
di: Aboeitta, Ahmed, et al.
Pubblicazione: (2025)
Unifying Speech Recognition, Synthesis and Conversion with Autoregressive Transformers
di: Cai, Runyuan, et al.
Pubblicazione: (2026)
di: Cai, Runyuan, et al.
Pubblicazione: (2026)
Transfer Learning-Based Deep Residual Learning for Speech Recognition in Clean and Noisy Environments
di: Djeffal, Noussaiba, et al.
Pubblicazione: (2025)
di: Djeffal, Noussaiba, et al.
Pubblicazione: (2025)
Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement
di: Wang, Junyu, et al.
Pubblicazione: (2024)
di: Wang, Junyu, et al.
Pubblicazione: (2024)
Unsupervised Speech Enhancement using Data-defined Priors
di: Klement, Dominik, et al.
Pubblicazione: (2025)
di: Klement, Dominik, et al.
Pubblicazione: (2025)
Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues
di: Hussain, Tassadaq, et al.
Pubblicazione: (2024)
di: Hussain, Tassadaq, et al.
Pubblicazione: (2024)
Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation
di: Ruggiero, Giuseppe, et al.
Pubblicazione: (2025)
di: Ruggiero, Giuseppe, et al.
Pubblicazione: (2025)
Enhancing Speech Emotion Recognition through Segmental Average Pooling of Self-Supervised Learning Features
di: Hyeon, Jonghwan, et al.
Pubblicazione: (2024)
di: Hyeon, Jonghwan, et al.
Pubblicazione: (2024)
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
di: Lin, Zijian, et al.
Pubblicazione: (2025)
di: Lin, Zijian, et al.
Pubblicazione: (2025)
Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
di: Bae, Hanbin, et al.
Pubblicazione: (2024)
di: Bae, Hanbin, et al.
Pubblicazione: (2024)
Speech Quality Assessment Model Based on Mixture of Experts: System-Level Performance Enhancement and Utterance-Level Challenge Analysis
di: Hu, Xintong, et al.
Pubblicazione: (2025)
di: Hu, Xintong, et al.
Pubblicazione: (2025)
BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching Models
di: Liang, Susan, et al.
Pubblicazione: (2025)
di: Liang, Susan, et al.
Pubblicazione: (2025)
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
di: Wang, Helin, et al.
Pubblicazione: (2025)
di: Wang, Helin, et al.
Pubblicazione: (2025)
Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models
di: Chen, Sijing, et al.
Pubblicazione: (2024)
di: Chen, Sijing, et al.
Pubblicazione: (2024)
Speech Enhancement Based on Drifting Models
di: Xu, Liang, et al.
Pubblicazione: (2026)
di: Xu, Liang, et al.
Pubblicazione: (2026)
A Mel Spectrogram Enhancement Paradigm Based on CWT in Speech Synthesis
di: Hu, Guoqiang, et al.
Pubblicazione: (2024)
di: Hu, Guoqiang, et al.
Pubblicazione: (2024)
Continuous Modeling of the Denoising Process for Speech Enhancement Based on Deep Learning
di: Guo, Zilu, et al.
Pubblicazione: (2023)
di: Guo, Zilu, et al.
Pubblicazione: (2023)
xLSTM-SENet: xLSTM for Single-Channel Speech Enhancement
di: Kühne, Nikolai Lund, et al.
Pubblicazione: (2025)
di: Kühne, Nikolai Lund, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Universal Speech Enhancement with Regression and Generative Mamba
di: Chao, Rong, et al.
Pubblicazione: (2025) -
An Investigation of Incorporating Mamba for Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2024) -
Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
di: Khan, Muhammad Salman, et al.
Pubblicazione: (2024) -
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2025) -
CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR
di: Shao, Nian, et al.
Pubblicazione: (2025)