FlowDec: A flow-based full-band general audio codec with high perceptual quality
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Welker, Simon, Le, Matthew, Chen, Ricky T. Q., Hsu, Wei-Ning, Gerkmann, Timo, Richard, Alexander, Wu, Yi-Chiao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Real-Time Streaming Mel Vocoding with Generative Flow Matching
von: Welker, Simon, et al.
Veröffentlicht: (2025)
von: Welker, Simon, et al.
Veröffentlicht: (2025)
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Speaker anonymization using neural audio codec language models
von: Panariello, Michele, et al.
Veröffentlicht: (2023)
von: Panariello, Michele, et al.
Veröffentlicht: (2023)
Modeling strategies for speech enhancement in the latent space of a neural audio codec
von: Kammoun, Sofiene, et al.
Veröffentlicht: (2025)
von: Kammoun, Sofiene, et al.
Veröffentlicht: (2025)
EMOCONV-DIFF: Diffusion-based Speech Emotion Conversion for Non-parallel and In-the-wild Data
von: Prabhu, Navin Raj, et al.
Veröffentlicht: (2023)
von: Prabhu, Navin Raj, et al.
Veröffentlicht: (2023)
StoRM: A Diffusion-based Stochastic Regeneration Model for Speech Enhancement and Dereverberation
von: Lemercier, Jean-Marie, et al.
Veröffentlicht: (2022)
von: Lemercier, Jean-Marie, et al.
Veröffentlicht: (2022)
The PESQetarian: On the Relevance of Goodhart's Law for Speech Enhancement
von: de Oliveira, Danilo, et al.
Veröffentlicht: (2024)
von: de Oliveira, Danilo, et al.
Veröffentlicht: (2024)
Diffusion Buffer for Online Generative Speech Enhancement
von: Lay, Bunlong, et al.
Veröffentlicht: (2025)
von: Lay, Bunlong, et al.
Veröffentlicht: (2025)
ReverbFX: A Dataset of Room Impulse Responses Derived from Reverb Effect Plugins for Singing Voice Dereverberation
von: Richter, Julius, et al.
Veröffentlicht: (2025)
von: Richter, Julius, et al.
Veröffentlicht: (2025)
EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation
von: Richter, Julius, et al.
Veröffentlicht: (2024)
von: Richter, Julius, et al.
Veröffentlicht: (2024)
BUDDy: Single-Channel Blind Unsupervised Dereverberation with Diffusion Models
von: Moliner, Eloi, et al.
Veröffentlicht: (2024)
von: Moliner, Eloi, et al.
Veröffentlicht: (2024)
Unsupervised Blind Joint Dereverberation and Room Acoustics Estimation with Diffusion Models
von: Lemercier, Jean-Marie, et al.
Veröffentlicht: (2024)
von: Lemercier, Jean-Marie, et al.
Veröffentlicht: (2024)
Speech Enhancement and Dereverberation with Diffusion-based Generative Models
von: Richter, Julius, et al.
Veröffentlicht: (2022)
von: Richter, Julius, et al.
Veröffentlicht: (2022)
Investigating Training Objectives for Generative Speech Enhancement
von: Richter, Julius, et al.
Veröffentlicht: (2024)
von: Richter, Julius, et al.
Veröffentlicht: (2024)
Do We Need EMA for Diffusion-Based Speech Enhancement? Toward a Magnitude-Preserving Network Architecture
von: Richter, Julius, et al.
Veröffentlicht: (2025)
von: Richter, Julius, et al.
Veröffentlicht: (2025)
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
von: Pepino, Leonardo, et al.
Veröffentlicht: (2023)
von: Pepino, Leonardo, et al.
Veröffentlicht: (2023)
Non-intrusive Speech Quality Assessment with Diffusion Models Trained on Clean Speech
von: de Oliveira, Danilo, et al.
Veröffentlicht: (2024)
von: de Oliveira, Danilo, et al.
Veröffentlicht: (2024)
Multi-channel Speech Separation Using Spatially Selective Deep Non-linear Filters
von: Tesch, Kristina, et al.
Veröffentlicht: (2023)
von: Tesch, Kristina, et al.
Veröffentlicht: (2023)
An Analysis of the Variance of Diffusion-based Speech Enhancement
von: Lay, Bunlong, et al.
Veröffentlicht: (2024)
von: Lay, Bunlong, et al.
Veröffentlicht: (2024)
Steering Deep Non-Linear Spatially Selective Filters for Weakly Guided Extraction of Moving Speakers in Dynamic Scenarios
von: Kienegger, Jakob, et al.
Veröffentlicht: (2025)
von: Kienegger, Jakob, et al.
Veröffentlicht: (2025)
Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
von: Kienegger, Jakob, et al.
Veröffentlicht: (2026)
von: Kienegger, Jakob, et al.
Veröffentlicht: (2026)
Adaptive Rotary Steering with Joint Autoregression for Robust Extraction of Closely Moving Speakers in Dynamic Scenarios
von: Kienegger, Jakob, et al.
Veröffentlicht: (2026)
von: Kienegger, Jakob, et al.
Veröffentlicht: (2026)
Diffusion Models for Audio Restoration
von: Lemercier, Jean-Marie, et al.
Veröffentlicht: (2024)
von: Lemercier, Jean-Marie, et al.
Veröffentlicht: (2024)
LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models
von: Richter, Julius, et al.
Veröffentlicht: (2025)
von: Richter, Julius, et al.
Veröffentlicht: (2025)
Are These Even Words? Quantifying the Gibberishness of Generative Speech Models
von: de Oliveira, Danilo, et al.
Veröffentlicht: (2025)
von: de Oliveira, Danilo, et al.
Veröffentlicht: (2025)
Scaling up masked audio encoder learning for general audio classification
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2024)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2024)
ComplexDec: A Domain-robust High-fidelity Neural Audio Codec with Complex Spectrum Modeling
von: Wu, Yi-Chiao, et al.
Veröffentlicht: (2025)
von: Wu, Yi-Chiao, et al.
Veröffentlicht: (2025)
From the perspective of perceptual speech quality: The robustness of frequency bands to noise
von: Fan, Junyi, et al.
Veröffentlicht: (2025)
von: Fan, Junyi, et al.
Veröffentlicht: (2025)
FreeCodec: A disentangled neural speech codec with fewer tokens
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
A tunable binaural audio telepresence system capable of balancing immersive and enhanced modes
von: Hsu, Yicheng, et al.
Veröffentlicht: (2024)
von: Hsu, Yicheng, et al.
Veröffentlicht: (2024)
Are Modern Speech Enhancement Systems Vulnerable to Adversarial Attacks?
von: Makarov, Rostislav, et al.
Veröffentlicht: (2025)
von: Makarov, Rostislav, et al.
Veröffentlicht: (2025)
Mask-Weighted Spatial Likelihood Coding for Speaker-Independent Joint Localization and Mask Estimation
von: Kienegger, Jakob, et al.
Veröffentlicht: (2024)
von: Kienegger, Jakob, et al.
Veröffentlicht: (2024)
Bone-conduction Guided Multimodal Speech Enhancement with Conditional Diffusion Models
von: Khanagha, Sina, et al.
Veröffentlicht: (2026)
von: Khanagha, Sina, et al.
Veröffentlicht: (2026)
MBCodec:Thorough disentangle for high-fidelity audio compression
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
von: Yang, Mu, et al.
Veröffentlicht: (2024)
von: Yang, Mu, et al.
Veröffentlicht: (2024)
ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
von: Jing, Xin, et al.
Veröffentlicht: (2024)
von: Jing, Xin, et al.
Veröffentlicht: (2024)
Self-Steering Deep Non-Linear Spatially Selective Filters for Efficient Extraction of Moving Speakers under Weak Guidance
von: Kienegger, Jakob, et al.
Veröffentlicht: (2025)
von: Kienegger, Jakob, et al.
Veröffentlicht: (2025)
Self-supervised learning method using multiple sampling strategies for general-purpose audio representation
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)
Wind Noise Reduction with a Diffusion-based Stochastic Regeneration Model
von: Lemercier, Jean-Marie, et al.
Veröffentlicht: (2023)
von: Lemercier, Jean-Marie, et al.
Veröffentlicht: (2023)
Towards audio language modeling -- an overview
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Real-Time Streaming Mel Vocoding with Generative Flow Matching
von: Welker, Simon, et al.
Veröffentlicht: (2025) -
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
von: Wu, Haibin, et al.
Veröffentlicht: (2024) -
Speaker anonymization using neural audio codec language models
von: Panariello, Michele, et al.
Veröffentlicht: (2023) -
Modeling strategies for speech enhancement in the latent space of a neural audio codec
von: Kammoun, Sofiene, et al.
Veröffentlicht: (2025) -
EMOCONV-DIFF: Diffusion-based Speech Emotion Conversion for Non-parallel and In-the-wild Data
von: Prabhu, Navin Raj, et al.
Veröffentlicht: (2023)