Linguistic Knowledge Transfer Learning for Speech Enhancement
Fuente:
arXiv
Salvato in:
| Autori principali: | Hung, Kuo-Hsuan, Lu, Xugang, Fu, Szu-Wei, Tseng, Huan-Hsin, Lin, Hsin-Yi, Lin, Chii-Wann, Tsao, Yu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024)
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024)
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
di: Ren, Wenze, et al.
Pubblicazione: (2024)
di: Ren, Wenze, et al.
Pubblicazione: (2024)
Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
di: Khan, Muhammad Salman, et al.
Pubblicazione: (2024)
di: Khan, Muhammad Salman, et al.
Pubblicazione: (2024)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2025)
di: Chao, Rong, et al.
Pubblicazione: (2025)
Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
di: Ren, Wenze, et al.
Pubblicazione: (2024)
di: Ren, Wenze, et al.
Pubblicazione: (2024)
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement
di: Lin, Meng-Ping, et al.
Pubblicazione: (2025)
di: Lin, Meng-Ping, et al.
Pubblicazione: (2025)
A Study on Incorporating Whisper for Robust Speech Assessment
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
Universal Speech Enhancement with Regression and Generative Mamba
di: Chao, Rong, et al.
Pubblicazione: (2025)
di: Chao, Rong, et al.
Pubblicazione: (2025)
HighRateMOS: Sampling-Rate Aware Modeling for Speech Quality Assessment
di: Ren, Wenze, et al.
Pubblicazione: (2025)
di: Ren, Wenze, et al.
Pubblicazione: (2025)
Cross-modal Knowledge Transfer Learning as Graph Matching Based on Optimal Transport for ASR
di: Lu, Xugang, et al.
Pubblicazione: (2025)
di: Lu, Xugang, et al.
Pubblicazione: (2025)
EffortNet: A Deep Learning Framework for Objective Assessment of Speech Enhancement Technologies Using EEG-Based Alpha Oscillations
di: Sung, Ching-Chih, et al.
Pubblicazione: (2025)
di: Sung, Ching-Chih, et al.
Pubblicazione: (2025)
Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2026)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2026)
The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
di: Huang, Wen-Chin, et al.
Pubblicazione: (2024)
di: Huang, Wen-Chin, et al.
Pubblicazione: (2024)
Deep Learning-based Non-Intrusive Multi-Objective Speech Assessment Model with Cross-Domain Features
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2021)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2021)
MOS-Bias: From Hidden Gender Bias to Gender-Aware Speech Quality Assessment
di: Ren, Wenze, et al.
Pubblicazione: (2026)
di: Ren, Wenze, et al.
Pubblicazione: (2026)
Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR
di: Lu, Xugang, et al.
Pubblicazione: (2024)
di: Lu, Xugang, et al.
Pubblicazione: (2024)
An Investigation of Incorporating Mamba for Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2024)
di: Chao, Rong, et al.
Pubblicazione: (2024)
LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
di: Chen, Chih-Ning, et al.
Pubblicazione: (2026)
di: Chen, Chih-Ning, et al.
Pubblicazione: (2026)
A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2024)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2024)
Combining Deterministic Enhanced Conditions with Dual-Streaming Encoding for Diffusion-Based Speech Enhancement
di: Shi, Hao, et al.
Pubblicazione: (2025)
di: Shi, Hao, et al.
Pubblicazione: (2025)
Speech Intelligibility Assessment with Uncertainty-Aware Whisper Embeddings and sLSTM
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
Restorative Speech Enhancement: A Progressive Approach Using SE and Codec Modules
di: Chiang, Hsin-Tien, et al.
Pubblicazione: (2024)
di: Chiang, Hsin-Tien, et al.
Pubblicazione: (2024)
From Evaluation to Optimization: Neural Speech Assessment for Downstream Applications
di: Tsao, Yu
Pubblicazione: (2025)
di: Tsao, Yu
Pubblicazione: (2025)
Leveraging Self-Supervised Audio-Visual Pretrained Models to Improve Vocoded Speech Intelligibility in Cochlear Implant Simulation
di: Lai, Richard Lee, et al.
Pubblicazione: (2023)
di: Lai, Richard Lee, et al.
Pubblicazione: (2023)
Feature Importance across Domains for Improving Non-Intrusive Speech Intelligibility Prediction in Hearing Aids
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
A Study on Zero-Shot Non-Intrusive Speech Intelligibility for Hearing Aids Using Large Language Models
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
MC-SEMamba: A Simple Multi-channel Extension of SEMamba
di: Ting, Wen-Yuan, et al.
Pubblicazione: (2024)
di: Ting, Wen-Yuan, et al.
Pubblicazione: (2024)
A Study on Speech Assessment with Visual Cues
di: Ahmed, Shafique, et al.
Pubblicazione: (2025)
di: Ahmed, Shafique, et al.
Pubblicazione: (2025)
Tracking Listener Attention: Gaze-Guided Audio-Visual Speech Enhancement Framework
di: Yang, Hsiang-Cheng, et al.
Pubblicazione: (2026)
di: Yang, Hsiang-Cheng, et al.
Pubblicazione: (2026)
PiCoGen2: Piano cover generation with transfer learning approach and weakly aligned data
di: Tan, Chih-Pin, et al.
Pubblicazione: (2024)
di: Tan, Chih-Pin, et al.
Pubblicazione: (2024)
SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models
di: Yin, Chun, et al.
Pubblicazione: (2024)
di: Yin, Chun, et al.
Pubblicazione: (2024)
An Investigation on Combining Geometry and Consistency Constraints into Phase Estimation for Speech Enhancement
di: Ho, Chun-Wei, et al.
Pubblicazione: (2025)
di: Ho, Chun-Wei, et al.
Pubblicazione: (2025)
Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues
di: Hussain, Tassadaq, et al.
Pubblicazione: (2024)
di: Hussain, Tassadaq, et al.
Pubblicazione: (2024)
Universal Robust Speech Adaptation for Cross-Domain Speech Recognition and Enhancement
di: Wang, Chien-Chun, et al.
Pubblicazione: (2026)
di: Wang, Chien-Chun, et al.
Pubblicazione: (2026)
Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech Edits
di: Huang, Sung-Feng, et al.
Pubblicazione: (2025)
di: Huang, Sung-Feng, et al.
Pubblicazione: (2025)
Improving the Robustness and Clinical Applicability of Automatic Respiratory Sound Classification Using Deep Learning-Based Audio Enhancement: Algorithm Development and Validation
di: Tzeng, Jing-Tong, et al.
Pubblicazione: (2024)
di: Tzeng, Jing-Tong, et al.
Pubblicazione: (2024)
DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
di: Lu, Ke-Han, et al.
Pubblicazione: (2024)
di: Lu, Ke-Han, et al.
Pubblicazione: (2024)
Integrated Multi-Level Knowledge Distillation for Enhanced Speaker Verification
di: Yang, Wenhao, et al.
Pubblicazione: (2024)
di: Yang, Wenhao, et al.
Pubblicazione: (2024)
Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
CodecFake+: A Large-Scale Neural Audio Codec-Based Deepfake Speech Dataset
di: Chen, Xuanjun, et al.
Pubblicazione: (2025)
di: Chen, Xuanjun, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024) -
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
di: Ren, Wenze, et al.
Pubblicazione: (2024) -
Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
di: Khan, Muhammad Salman, et al.
Pubblicazione: (2024) -
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2025) -
Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
di: Ren, Wenze, et al.
Pubblicazione: (2024)