Gespeichert in:
| Hauptverfasser: | Li, Zixuan, Zhang, Xueliang, Zhao, Changjiang, Gao, Shuai, Miao, Lei, Yan, Zhipeng, Sun, Ying, Zhu, Chong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2511.05945 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Two-stage Audio-Visual Target Speaker Extraction System for Real-Time Processing On Edge Device
von: Li, Zixuan, et al.
Veröffentlicht: (2025)
von: Li, Zixuan, et al.
Veröffentlicht: (2025)
SICRN: Advancing Speech Enhancement through State Space Model and Inplace Convolution Techniques
von: Zhao, Changjiang, et al.
Veröffentlicht: (2024)
von: Zhao, Changjiang, et al.
Veröffentlicht: (2024)
Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
von: Khan, Muhammad Salman, et al.
Veröffentlicht: (2024)
von: Khan, Muhammad Salman, et al.
Veröffentlicht: (2024)
Speech Loudness in Broadcasting and Streaming
von: Torcoli, Matteo, et al.
Veröffentlicht: (2024)
von: Torcoli, Matteo, et al.
Veröffentlicht: (2024)
Attention-Based Beamformer For Multi-Channel Speech Enhancement
von: Bai, Jinglin, et al.
Veröffentlicht: (2024)
von: Bai, Jinglin, et al.
Veröffentlicht: (2024)
Hallucination in Perceptual Metric-Driven Speech Enhancement Networks
von: Close, George, et al.
Veröffentlicht: (2024)
von: Close, George, et al.
Veröffentlicht: (2024)
Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
von: Sutherland, Robert, et al.
Veröffentlicht: (2024)
von: Sutherland, Robert, et al.
Veröffentlicht: (2024)
An Explicit Consistency-Preserving Loss Function for Phase Reconstruction and Speech Enhancement
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
URGENT-PK: Perceptually-Aligned Ranking Model Designed for Speech Enhancement Competition
von: Wang, Jiahe, et al.
Veröffentlicht: (2025)
von: Wang, Jiahe, et al.
Veröffentlicht: (2025)
Robust Target Speaker Direction of Arrival Estimation
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms
von: Li, Sirui, et al.
Veröffentlicht: (2025)
von: Li, Sirui, et al.
Veröffentlicht: (2025)
Attention-Enhanced Short-Time Wiener Solution for Acoustic Echo Cancellation
von: Zhao, Fei, et al.
Veröffentlicht: (2024)
von: Zhao, Fei, et al.
Veröffentlicht: (2024)
NOMAD: Unsupervised Learning of Perceptual Embeddings for Speech Enhancement and Non-matching Reference Audio Quality Assessment
von: Ragano, Alessandro, et al.
Veröffentlicht: (2023)
von: Ragano, Alessandro, et al.
Veröffentlicht: (2023)
Conditional Latent Diffusion-Based Speech Enhancement Via Dual Context Learning
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
From Coarse to Fine: Recursive Audio-Visual Semantic Enhancement for Speech Separation
von: Xue, Ke, et al.
Veröffentlicht: (2025)
von: Xue, Ke, et al.
Veröffentlicht: (2025)
On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation
von: Cheng, Changhao, et al.
Veröffentlicht: (2026)
von: Cheng, Changhao, et al.
Veröffentlicht: (2026)
A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
von: Wang, Jingyuan, et al.
Veröffentlicht: (2024)
von: Wang, Jingyuan, et al.
Veröffentlicht: (2024)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
Adaptive Convolution for CNN-based Speech Enhancement Models
von: Wang, Dahan, et al.
Veröffentlicht: (2025)
von: Wang, Dahan, et al.
Veröffentlicht: (2025)
Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network
von: Chen, Yanan, et al.
Veröffentlicht: (2024)
von: Chen, Yanan, et al.
Veröffentlicht: (2024)
Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses
von: Zhao, Shengkui, et al.
Veröffentlicht: (2021)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2021)
Decoders Laugh as Loud as Encoders
von: Borodach, Eli, et al.
Veröffentlicht: (2025)
von: Borodach, Eli, et al.
Veröffentlicht: (2025)
Vector Quantized Diffusion Model Based Speech Bandwidth Extension
von: Fang, Yuan, et al.
Veröffentlicht: (2024)
von: Fang, Yuan, et al.
Veröffentlicht: (2024)
VibOmni: Towards Scalable Bone-conduction Speech Enhancement on Earables
von: He, Lixing, et al.
Veröffentlicht: (2025)
von: He, Lixing, et al.
Veröffentlicht: (2025)
Speech Synthesis along Perceptual Voice Quality Dimensions
von: Rautenberg, Frederik, et al.
Veröffentlicht: (2025)
von: Rautenberg, Frederik, et al.
Veröffentlicht: (2025)
Low-latency Speech Enhancement via Speech Token Generation
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
Universal Discrete-Domain Speech Enhancement
von: Liu, Fei, et al.
Veröffentlicht: (2025)
von: Liu, Fei, et al.
Veröffentlicht: (2025)
Magnitude-Phase Dual-Path Speech Enhancement Network based on Self-Supervised Embedding and Perceptual Contrast Stretch Boosting
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2025)
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2025)
Multi-Channel Speech Enhancement for Cocktail Party Speech Emotion Recognition
von: Chen, Youjun, et al.
Veröffentlicht: (2026)
von: Chen, Youjun, et al.
Veröffentlicht: (2026)
Frequency-Weighted Training Losses for Phoneme-Level DNN-based Speech Enhancement
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2025)
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2025)
BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
Diffusion-based Frameworks for Unsupervised Speech Enhancement
von: Ayilo, Jean-Eudes, et al.
Veröffentlicht: (2026)
von: Ayilo, Jean-Eudes, et al.
Veröffentlicht: (2026)
The Interstellar Scintillation of the Radio-Loud Magnetar XTE J1810-197
von: Wang, Rui, et al.
Veröffentlicht: (2026)
von: Wang, Rui, et al.
Veröffentlicht: (2026)
Study of Lightweight Transformer Architectures for Single-Channel Speech Enhancement
von: Zhao, Haixin, et al.
Veröffentlicht: (2025)
von: Zhao, Haixin, et al.
Veröffentlicht: (2025)
Complex-Cycle-Consistent Diffusion Model for Monaural Speech Enhancement
von: Li, Yi, et al.
Veröffentlicht: (2024)
von: Li, Yi, et al.
Veröffentlicht: (2024)
Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification
von: Zhao, Yiyang, et al.
Veröffentlicht: (2025)
von: Zhao, Yiyang, et al.
Veröffentlicht: (2025)
Critique-out-Loud Reward Models
von: Ankner, Zachary, et al.
Veröffentlicht: (2024)
von: Ankner, Zachary, et al.
Veröffentlicht: (2024)
Recite Your Ask Out Loud
Veröffentlicht: (2025)
Veröffentlicht: (2025)
Robust One-step Speech Enhancement via Consistency Distillation
von: Xu, Liang, et al.
Veröffentlicht: (2025)
von: Xu, Liang, et al.
Veröffentlicht: (2025)
A Probabilistic Generative Model for Spectral Speech Enhancement
von: Hidalgo-Araya, Marco, et al.
Veröffentlicht: (2026)
von: Hidalgo-Araya, Marco, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Two-stage Audio-Visual Target Speaker Extraction System for Real-Time Processing On Edge Device
von: Li, Zixuan, et al.
Veröffentlicht: (2025) -
SICRN: Advancing Speech Enhancement through State Space Model and Inplace Convolution Techniques
von: Zhao, Changjiang, et al.
Veröffentlicht: (2024) -
Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
von: Khan, Muhammad Salman, et al.
Veröffentlicht: (2024) -
Speech Loudness in Broadcasting and Streaming
von: Torcoli, Matteo, et al.
Veröffentlicht: (2024) -
Attention-Based Beamformer For Multi-Channel Speech Enhancement
von: Bai, Jinglin, et al.
Veröffentlicht: (2024)