Rethinking Training Targets, Architectures and Data Quality for Universal Speech Enhancement
Fuente:
arXiv
Salvato in:
| Autori principali: | Fu, Szu-Wei, Chao, Rong, Yang, Xuesong, Huang, Sung-Feng, Zezario, Ryandhimas E., Nasretdinov, Rauf, Jukić, Ante, Tsao, Yu, Wang, Yu-Chiang Frank |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Universal Speech Enhancement with Regression and Generative Mamba
di: Chao, Rong, et al.
Pubblicazione: (2025)
di: Chao, Rong, et al.
Pubblicazione: (2025)
Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement
di: Nasretdinov, Rauf, et al.
Pubblicazione: (2025)
di: Nasretdinov, Rauf, et al.
Pubblicazione: (2025)
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024)
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024)
A Study on Incorporating Whisper for Robust Speech Assessment
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2026)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2026)
The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
di: Huang, Wen-Chin, et al.
Pubblicazione: (2024)
di: Huang, Wen-Chin, et al.
Pubblicazione: (2024)
Deep Learning-based Non-Intrusive Multi-Objective Speech Assessment Model with Cross-Domain Features
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2021)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2021)
A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2024)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2024)
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
di: Ku, Pin-Jui, et al.
Pubblicazione: (2024)
di: Ku, Pin-Jui, et al.
Pubblicazione: (2024)
Non-Intrusive Intelligibility Prediction for Hearing Aids: Recent Advances, Trends, and Challenges
di: Zezario, Ryandhimas E.
Pubblicazione: (2025)
di: Zezario, Ryandhimas E.
Pubblicazione: (2025)
Speech Intelligibility Assessment with Uncertainty-Aware Whisper Embeddings and sLSTM
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
Multi-Task Pseudo-Label Learning for Non-Intrusive Speech Quality Assessment Model
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
Feature Importance across Domains for Improving Non-Intrusive Speech Intelligibility Prediction in Hearing Aids
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
A Study on Zero-Shot Non-Intrusive Speech Intelligibility for Hearing Aids Using Large Language Models
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
A Study on Speech Assessment with Visual Cues
di: Ahmed, Shafique, et al.
Pubblicazione: (2025)
di: Ahmed, Shafique, et al.
Pubblicazione: (2025)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2025)
di: Chao, Rong, et al.
Pubblicazione: (2025)
An Investigation of Incorporating Mamba for Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2024)
di: Chao, Rong, et al.
Pubblicazione: (2024)
STSM-FiLM: A FiLM-Conditioned Neural Architecture for Time-Scale Modification of Speech
di: Wisnu, Dyah A. M. G., et al.
Pubblicazione: (2025)
di: Wisnu, Dyah A. M. G., et al.
Pubblicazione: (2025)
HAAQI-Net: A Non-intrusive Neural Music Audio Quality Assessment Model for Hearing Aids
di: Wisnu, Dyah A. M. G., et al.
Pubblicazione: (2024)
di: Wisnu, Dyah A. M. G., et al.
Pubblicazione: (2024)
Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech Edits
di: Huang, Sung-Feng, et al.
Pubblicazione: (2025)
di: Huang, Sung-Feng, et al.
Pubblicazione: (2025)
Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings
di: Wisnu, Dyah A. M. G., et al.
Pubblicazione: (2025)
di: Wisnu, Dyah A. M. G., et al.
Pubblicazione: (2025)
NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
di: Casanova, Edresson, et al.
Pubblicazione: (2025)
di: Casanova, Edresson, et al.
Pubblicazione: (2025)
Neuro-MSBG: An End-to-End Neural Model for Hearing Loss Simulation
di: Yuan, Hui-Guan, et al.
Pubblicazione: (2025)
di: Yuan, Hui-Guan, et al.
Pubblicazione: (2025)
Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
di: Khan, Muhammad Salman, et al.
Pubblicazione: (2024)
di: Khan, Muhammad Salman, et al.
Pubblicazione: (2024)
NeuroAMP: A Novel End-to-end General Purpose Deep Neural Amplifier for Personalized Hearing Aids
di: Ahmed, Shafique, et al.
Pubblicazione: (2025)
di: Ahmed, Shafique, et al.
Pubblicazione: (2025)
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
di: Ren, Wenze, et al.
Pubblicazione: (2024)
di: Ren, Wenze, et al.
Pubblicazione: (2024)
Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
di: Casanova, Edresson, et al.
Pubblicazione: (2024)
di: Casanova, Edresson, et al.
Pubblicazione: (2024)
Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues
di: Hussain, Tassadaq, et al.
Pubblicazione: (2024)
di: Hussain, Tassadaq, et al.
Pubblicazione: (2024)
DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
di: Lu, Ke-Han, et al.
Pubblicazione: (2024)
di: Lu, Ke-Han, et al.
Pubblicazione: (2024)
Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
di: Ren, Wenze, et al.
Pubblicazione: (2024)
di: Ren, Wenze, et al.
Pubblicazione: (2024)
RankUp: Boosting Semi-Supervised Regression with an Auxiliary Ranking Classifier
di: Huang, Pin-Yen, et al.
Pubblicazione: (2024)
di: Huang, Pin-Yen, et al.
Pubblicazione: (2024)
Restorative Speech Enhancement: A Progressive Approach Using SE and Codec Modules
di: Chiang, Hsin-Tien, et al.
Pubblicazione: (2024)
di: Chiang, Hsin-Tien, et al.
Pubblicazione: (2024)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
di: Wang, Kuan-Chen, et al.
Pubblicazione: (2024)
di: Wang, Kuan-Chen, et al.
Pubblicazione: (2024)
Universal Discrete-Domain Speech Enhancement
di: Liu, Fei, et al.
Pubblicazione: (2025)
di: Liu, Fei, et al.
Pubblicazione: (2025)
Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network
di: Chen, Yanan, et al.
Pubblicazione: (2024)
di: Chen, Yanan, et al.
Pubblicazione: (2024)
QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
di: Wang, Siyin, et al.
Pubblicazione: (2025)
di: Wang, Siyin, et al.
Pubblicazione: (2025)
LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
di: Chen, Chih-Ning, et al.
Pubblicazione: (2026)
di: Chen, Chih-Ning, et al.
Pubblicazione: (2026)
GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement
di: Rong, Xiaobin, et al.
Pubblicazione: (2026)
di: Rong, Xiaobin, et al.
Pubblicazione: (2026)
Towards Environmental Preference Based Speech Enhancement For Individualised Multi-Modal Hearing Aids
di: Kirton-Wingate, Jasper, et al.
Pubblicazione: (2024)
di: Kirton-Wingate, Jasper, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Universal Speech Enhancement with Regression and Generative Mamba
di: Chao, Rong, et al.
Pubblicazione: (2025) -
Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement
di: Nasretdinov, Rauf, et al.
Pubblicazione: (2025) -
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024) -
A Study on Incorporating Whisper for Robust Speech Assessment
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023) -
Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2026)