FlowMAC: Conditional Flow Matching for Audio Coding at Low Bit Rates
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pia, Nicola, Strauss, Martin, Multrus, Markus, Edler, Bernd |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Neural Speech Coding for Real-time Communications using Constant Bitrate Scalar Quantization
von: Brendel, Andreas, et al.
Veröffentlicht: (2024)
von: Brendel, Andreas, et al.
Veröffentlicht: (2024)
Parametric Object Coding in IVAS: Efficient Coding of Multiple Audio Objects at Low Bit Rates
von: Eichenseer, Andrea, et al.
Veröffentlicht: (2025)
von: Eichenseer, Andrea, et al.
Veröffentlicht: (2025)
A DNN Based Post-Filter to Enhance the Quality of Coded Speech in MDCT Domain
von: Gupta, Kishan, et al.
Veröffentlicht: (2022)
von: Gupta, Kishan, et al.
Veröffentlicht: (2022)
FlowTSE: Target Speaker Extraction with Flow Matching
von: Navon, Aviv, et al.
Veröffentlicht: (2025)
von: Navon, Aviv, et al.
Veröffentlicht: (2025)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
On Improving Error Resilience of Neural End-to-End Speech Coders
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
Drax: Speech Recognition with Discrete Flow Matching
von: Navon, Aviv, et al.
Veröffentlicht: (2025)
von: Navon, Aviv, et al.
Veröffentlicht: (2025)
Gaussian Flow Bridges for Audio Domain Transfer with Unpaired Data
von: Moliner, Eloi, et al.
Veröffentlicht: (2024)
von: Moliner, Eloi, et al.
Veröffentlicht: (2024)
Audio Match Cutting: Finding and Creating Matching Audio Transitions in Movies and Videos
von: Fedorishin, Dennis, et al.
Veröffentlicht: (2024)
von: Fedorishin, Dennis, et al.
Veröffentlicht: (2024)
RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2026)
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2026)
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
Synthesizer Sound Matching Using Audio Spectrogram Transformers
von: Bruford, Fred, et al.
Veröffentlicht: (2024)
von: Bruford, Fred, et al.
Veröffentlicht: (2024)
Fast Timing-Conditioned Latent Audio Diffusion
von: Evans, Zach, et al.
Veröffentlicht: (2024)
von: Evans, Zach, et al.
Veröffentlicht: (2024)
Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models
von: Feng, Chen, et al.
Veröffentlicht: (2025)
von: Feng, Chen, et al.
Veröffentlicht: (2025)
Real-Time Streaming Mel Vocoding with Generative Flow Matching
von: Welker, Simon, et al.
Veröffentlicht: (2025)
von: Welker, Simon, et al.
Veröffentlicht: (2025)
Conditional Generative Data Augmentation for Clinical Audio Datasets
von: Seibold, Matthias, et al.
Veröffentlicht: (2022)
von: Seibold, Matthias, et al.
Veröffentlicht: (2022)
DFADD: The Diffusion and Flow-Matching Based Audio Deepfake Dataset
von: Du, Jiawei, et al.
Veröffentlicht: (2024)
von: Du, Jiawei, et al.
Veröffentlicht: (2024)
Variable Bitrate Residual Vector Quantization for Audio Coding
von: Chae, Yunkee, et al.
Veröffentlicht: (2024)
von: Chae, Yunkee, et al.
Veröffentlicht: (2024)
LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
von: Guan, Wenhao, et al.
Veröffentlicht: (2024)
von: Guan, Wenhao, et al.
Veröffentlicht: (2024)
Generative Pre-training for Speech with Flow Matching
von: Liu, Alexander H., et al.
Veröffentlicht: (2023)
von: Liu, Alexander H., et al.
Veröffentlicht: (2023)
DeCoR: Defy Knowledge Forgetting by Predicting Earlier Audio Codes
von: Jiang, Xilin, et al.
Veröffentlicht: (2023)
von: Jiang, Xilin, et al.
Veröffentlicht: (2023)
Audio Classification of Low Feature Spectrograms Utilizing Convolutional Neural Networks
von: Elias, Noel
Veröffentlicht: (2024)
von: Elias, Noel
Veröffentlicht: (2024)
Mixture of Low-Rank Adapter Experts in Generalizable Audio Deepfake Detection
von: Laakkonen, Janne, et al.
Veröffentlicht: (2025)
von: Laakkonen, Janne, et al.
Veröffentlicht: (2025)
Intelligent Fault Diagnosis of Type and Severity in Low-Frequency, Low Bit-Depth Signals
von: Spadini, Tito, et al.
Veröffentlicht: (2024)
von: Spadini, Tito, et al.
Veröffentlicht: (2024)
Instabilities in Convnets for Raw Audio
von: Haider, Daniel, et al.
Veröffentlicht: (2023)
von: Haider, Daniel, et al.
Veröffentlicht: (2023)
EuleroDec: A Complex-Valued RVQ-VAE for Efficient and Robust Audio Coding
von: Cerovaz, Luca, et al.
Veröffentlicht: (2026)
von: Cerovaz, Luca, et al.
Veröffentlicht: (2026)
Self-Learning for Personalized Keyword Spotting on Ultra-Low-Power Audio Sensors
von: Rusci, Manuele, et al.
Veröffentlicht: (2024)
von: Rusci, Manuele, et al.
Veröffentlicht: (2024)
PTQ4ADM: Post-Training Quantization for Efficient Text Conditional Audio Diffusion Models
von: Vora, Jayneel, et al.
Veröffentlicht: (2024)
von: Vora, Jayneel, et al.
Veröffentlicht: (2024)
MusFlow: Multimodal Music Generation via Conditional Flow Matching
von: Song, Jiahao, et al.
Veröffentlicht: (2025)
von: Song, Jiahao, et al.
Veröffentlicht: (2025)
Diffusion-Based Speech Enhancement in Matched and Mismatched Conditions Using a Heun-Based Sampler
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2023)
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2023)
Transformer Redesign for Late Fusion of Audio-Text Features on Ultra-Low-Power Edge Hardware
von: Mitsis, Stavros, et al.
Veröffentlicht: (2025)
von: Mitsis, Stavros, et al.
Veröffentlicht: (2025)
A2SB: Audio-to-Audio Schrodinger Bridges
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
von: Primus, Paul, et al.
Veröffentlicht: (2025)
von: Primus, Paul, et al.
Veröffentlicht: (2025)
The Rarity of Musical Audio Signals Within the Space of Possible Audio Generation
von: Collins, Nick
Veröffentlicht: (2024)
von: Collins, Nick
Veröffentlicht: (2024)
Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
von: Primus, Paul, et al.
Veröffentlicht: (2024)
von: Primus, Paul, et al.
Veröffentlicht: (2024)
Fusing Audio and Metadata Embeddings Improves Language-based Audio Retrieval
von: Primus, Paul, et al.
Veröffentlicht: (2024)
von: Primus, Paul, et al.
Veröffentlicht: (2024)
Source Separation by Flow Matching
von: Scheibler, Robin, et al.
Veröffentlicht: (2025)
von: Scheibler, Robin, et al.
Veröffentlicht: (2025)
Compose Yourself: Average-Velocity Flow Matching for One-Step Speech Enhancement
von: Yang, Gang, et al.
Veröffentlicht: (2025)
von: Yang, Gang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Neural Speech Coding for Real-time Communications using Constant Bitrate Scalar Quantization
von: Brendel, Andreas, et al.
Veröffentlicht: (2024) -
Parametric Object Coding in IVAS: Efficient Coding of Multiple Audio Objects at Low Bit Rates
von: Eichenseer, Andrea, et al.
Veröffentlicht: (2025) -
A DNN Based Post-Filter to Enhance the Quality of Coded Speech in MDCT Domain
von: Gupta, Kishan, et al.
Veröffentlicht: (2022) -
FlowTSE: Target Speaker Extraction with Flow Matching
von: Navon, Aviv, et al.
Veröffentlicht: (2025) -
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)