PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rong, Xiaobin, Hu, Qinwen, Yesilbursa, Mansur, Wojcicki, Kamil, Lu, Jing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
Adapting WavLM for Speech Emotion Recognition
von: Diatlova, Daria, et al.
Veröffentlicht: (2024)
von: Diatlova, Daria, et al.
Veröffentlicht: (2024)
WavLM model ensemble for audio deepfake detection
von: Combei, David, et al.
Veröffentlicht: (2024)
von: Combei, David, et al.
Veröffentlicht: (2024)
SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
von: Lin, Jingru, et al.
Veröffentlicht: (2024)
von: Lin, Jingru, et al.
Veröffentlicht: (2024)
Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection
von: Stourbe, Theophile, et al.
Veröffentlicht: (2024)
von: Stourbe, Theophile, et al.
Veröffentlicht: (2024)
Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions
von: Mack, Wolfgang, et al.
Veröffentlicht: (2025)
von: Mack, Wolfgang, et al.
Veröffentlicht: (2025)
Audio Deepfake Detection with Self-Supervised WavLM and Multi-Fusion Attentive Classifier
von: Guo, Yinlin, et al.
Veröffentlicht: (2023)
von: Guo, Yinlin, et al.
Veröffentlicht: (2023)
Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech
von: Nakazawa, Kazushi
Veröffentlicht: (2026)
von: Nakazawa, Kazushi
Veröffentlicht: (2026)
Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation
von: Wang, Zheng, et al.
Veröffentlicht: (2026)
von: Wang, Zheng, et al.
Veröffentlicht: (2026)
TS-URGENet: A Three-stage Universal Robust and Generalizable Speech Enhancement Network
von: Rong, Xiaobin, et al.
Veröffentlicht: (2025)
von: Rong, Xiaobin, et al.
Veröffentlicht: (2025)
Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM
von: Sun, Zhaokai, et al.
Veröffentlicht: (2025)
von: Sun, Zhaokai, et al.
Veröffentlicht: (2025)
XWSB: A Blend System Utilizing XLS-R and WavLM with SLS Classifier detection system for SVDD 2024 Challenge
von: Zhang, Qishan, et al.
Veröffentlicht: (2024)
von: Zhang, Qishan, et al.
Veröffentlicht: (2024)
SNR-Progressive Model with Harmonic Compensation for Low-SNR Speech Enhancement
von: Hou, Zhongshu, et al.
Veröffentlicht: (2024)
von: Hou, Zhongshu, et al.
Veröffentlicht: (2024)
Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2025)
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2025)
Crowdsourced Multilingual Speech Intelligibility Testing
von: Lechler, Laura, et al.
Veröffentlicht: (2024)
von: Lechler, Laura, et al.
Veröffentlicht: (2024)
Low-Resource Audio Codec (LRAC): 2025 Challenge Description
von: Wojcicki, Kamil, et al.
Veröffentlicht: (2025)
von: Wojcicki, Kamil, et al.
Veröffentlicht: (2025)
GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
A Lightweight Hybrid Dual Channel Speech Enhancement System under Low-SNR Conditions
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
FNSE-SBGAN: Far-field Speech Enhancement with Schrodinger Bridge and Generative Adversarial Networks
von: Lei, Tong, et al.
Veröffentlicht: (2025)
von: Lei, Tong, et al.
Veröffentlicht: (2025)
Adaptive Convolution for CNN-based Speech Enhancement Models
von: Wang, Dahan, et al.
Veröffentlicht: (2025)
von: Wang, Dahan, et al.
Veröffentlicht: (2025)
Audio is all in one: speech-driven gesture synthetics using WavLM pre-trained model
von: Zhang, Fan, et al.
Veröffentlicht: (2023)
von: Zhang, Fan, et al.
Veröffentlicht: (2023)
UL-UNAS: Ultra-Lightweight U-Nets for Real-Time Speech Enhancement via Network Architecture Search
von: Rong, Xiaobin, et al.
Veröffentlicht: (2025)
von: Rong, Xiaobin, et al.
Veröffentlicht: (2025)
Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition
von: Kounadis-Bastian, Dionyssos, et al.
Veröffentlicht: (2024)
von: Kounadis-Bastian, Dionyssos, et al.
Veröffentlicht: (2024)
ProSE: Diffusion Priors for Speech Enhancement
von: Kumar, Sonal, et al.
Veröffentlicht: (2025)
von: Kumar, Sonal, et al.
Veröffentlicht: (2025)
Weakly Supervised Phonological Features for Pathological Speech Analysis
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2025)
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2025)
Low-latency Speech Enhancement via Speech Token Generation
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
Rethinking Flow and Diffusion Bridge Models for Speech Enhancement
von: Wang, Dahan, et al.
Veröffentlicht: (2026)
von: Wang, Dahan, et al.
Veröffentlicht: (2026)
Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabs
von: Pietroń, Marcin, et al.
Veröffentlicht: (2026)
von: Pietroń, Marcin, et al.
Veröffentlicht: (2026)
VoCodec: An Efficient Lightweight Low-Bitrate Speech Codec
von: Yang, Leyan, et al.
Veröffentlicht: (2026)
von: Yang, Leyan, et al.
Veröffentlicht: (2026)
Hallucination in Perceptual Metric-Driven Speech Enhancement Networks
von: Close, George, et al.
Veröffentlicht: (2024)
von: Close, George, et al.
Veröffentlicht: (2024)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
von: Chao, Rong, et al.
Veröffentlicht: (2025)
von: Chao, Rong, et al.
Veröffentlicht: (2025)
Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
Leveraging Discriminative Latent Representations for Conditioning GAN-Based Speech Enhancement
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2025)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2025)
WavCraft: Audio Editing and Generation with Large Language Models
von: Liang, Jinhua, et al.
Veröffentlicht: (2024)
von: Liang, Jinhua, et al.
Veröffentlicht: (2024)
CFMDCTCodec: A Low-Bitrate Neural Speech Codec with Noise-Prior-aware Conditional Flow Matching for MDCT-Spectral Enhancement
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2026)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2026)
Universal Speech Enhancement with Regression and Generative Mamba
von: Chao, Rong, et al.
Veröffentlicht: (2025)
von: Chao, Rong, et al.
Veröffentlicht: (2025)
Schrödinger Bridge for Generative Speech Enhancement
von: Jukić, Ante, et al.
Veröffentlicht: (2024)
von: Jukić, Ante, et al.
Veröffentlicht: (2024)
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection
von: Huang, Shangkun, et al.
Veröffentlicht: (2025)
von: Huang, Shangkun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026) -
UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026) -
Adapting WavLM for Speech Emotion Recognition
von: Diatlova, Daria, et al.
Veröffentlicht: (2024) -
WavLM model ensemble for audio deepfake detection
von: Combei, David, et al.
Veröffentlicht: (2024) -
SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
von: Lin, Jingru, et al.
Veröffentlicht: (2024)