CodecFlow: Efficient Bandwidth Extension via Conditional Flow Matching in Neural Codec Latent Space
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Bowen, Zhao, Junchuan, McLoughlin, Ian, Wang, Ye, Madhukumar, A S |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
Harmonic-Percussive Disentangled Neural Audio Codec for Bandwidth Extension
von: Giniès, Benoît, et al.
Veröffentlicht: (2025)
von: Giniès, Benoît, et al.
Veröffentlicht: (2025)
AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
von: Wang, Lu, et al.
Veröffentlicht: (2025)
von: Wang, Lu, et al.
Veröffentlicht: (2025)
FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
von: Ahasan, Md Mubtasim, et al.
Veröffentlicht: (2025)
von: Ahasan, Md Mubtasim, et al.
Veröffentlicht: (2025)
Neural Codec Source Tracing: Toward Comprehensive Attribution in Open-Set Condition
von: Xie, Yuankun, et al.
Veröffentlicht: (2025)
von: Xie, Yuankun, et al.
Veröffentlicht: (2025)
Drum Synthesis from Expressive Drum Grids via Neural Audio Codecs
von: Soiledis, Konstantinos, et al.
Veröffentlicht: (2026)
von: Soiledis, Konstantinos, et al.
Veröffentlicht: (2026)
PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders
von: Pan, Yu, et al.
Veröffentlicht: (2024)
von: Pan, Yu, et al.
Veröffentlicht: (2024)
The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs
von: Sadok, Samir, et al.
Veröffentlicht: (2026)
von: Sadok, Samir, et al.
Veröffentlicht: (2026)
HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling
von: Xue, Rongkun, et al.
Veröffentlicht: (2025)
von: Xue, Rongkun, et al.
Veröffentlicht: (2025)
Codec-Robust Attacks on Audio LLMs
von: Roh, Jaechul, et al.
Veröffentlicht: (2026)
von: Roh, Jaechul, et al.
Veröffentlicht: (2026)
SpectroStream: A Versatile Neural Codec for General Audio
von: Li, Yunpeng, et al.
Veröffentlicht: (2025)
von: Li, Yunpeng, et al.
Veröffentlicht: (2025)
Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
On the Design of Diffusion-based Neural Speech Codecs
von: Foti, Pietro, et al.
Veröffentlicht: (2025)
von: Foti, Pietro, et al.
Veröffentlicht: (2025)
Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
von: Zhang, Hanglei, et al.
Veröffentlicht: (2025)
von: Zhang, Hanglei, et al.
Veröffentlicht: (2025)
SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
Learning Source Disentanglement in Neural Audio Codec
von: Bie, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Bie, Xiaoyu, et al.
Veröffentlicht: (2024)
A General Model for Deepfake Speech Detection: Diverse Bonafide Resources or Diverse AI-Based Generators
von: Pham, Lam, et al.
Veröffentlicht: (2026)
von: Pham, Lam, et al.
Veröffentlicht: (2026)
Prototype based Masked Audio Model for Self-Supervised Learning of Sound Event Detection
von: Cai, Pengfei, et al.
Veröffentlicht: (2024)
von: Cai, Pengfei, et al.
Veröffentlicht: (2024)
MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection
von: Cai, Pengfei, et al.
Veröffentlicht: (2024)
von: Cai, Pengfei, et al.
Veröffentlicht: (2024)
MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation
von: Song, Yakun, et al.
Veröffentlicht: (2025)
von: Song, Yakun, et al.
Veröffentlicht: (2025)
Automated evaluation of children's speech fluency for low-resource languages
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
Flowing Straighter with Conditional Flow Matching for Accurate Speech Enhancement
von: Cross, Mattias, et al.
Veröffentlicht: (2025)
von: Cross, Mattias, et al.
Veröffentlicht: (2025)
Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt
von: Shi, Yanfeng, et al.
Veröffentlicht: (2026)
von: Shi, Yanfeng, et al.
Veröffentlicht: (2026)
Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection
von: Wu, Jinyang, et al.
Veröffentlicht: (2026)
von: Wu, Jinyang, et al.
Veröffentlicht: (2026)
DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers
von: Cao, Tianyu, et al.
Veröffentlicht: (2026)
von: Cao, Tianyu, et al.
Veröffentlicht: (2026)
Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
UniSRCodec: Unified and Low-Bitrate Single Codebook Codec with Sub-Band Reconstruction
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2026)
MIDI-VALLE: Improving Expressive Piano Performance Synthesis Through Neural Codec Language Modelling
von: Tang, Jingjing, et al.
Veröffentlicht: (2025)
von: Tang, Jingjing, et al.
Veröffentlicht: (2025)
Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries
von: Cai, Pengfei, et al.
Veröffentlicht: (2025)
von: Cai, Pengfei, et al.
Veröffentlicht: (2025)
Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems
von: Chen, Leduo, et al.
Veröffentlicht: (2026)
von: Chen, Leduo, et al.
Veröffentlicht: (2026)
AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling
von: Shi, Jiacheng, et al.
Veröffentlicht: (2026)
von: Shi, Jiacheng, et al.
Veröffentlicht: (2026)
Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
AlphaFlowTSE: One-Step Generative Target Speaker Extraction via Conditional AlphaFlow
von: Li, Duojia, et al.
Veröffentlicht: (2026)
von: Li, Duojia, et al.
Veröffentlicht: (2026)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents
von: Banerjee, Adhiraj, et al.
Veröffentlicht: (2025)
von: Banerjee, Adhiraj, et al.
Veröffentlicht: (2025)
Neural Speech and Audio Coding: Modern AI Technology Meets Traditional Codecs
von: Kim, Minje, et al.
Veröffentlicht: (2024)
von: Kim, Minje, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025) -
Harmonic-Percussive Disentangled Neural Audio Codec for Bandwidth Extension
von: Giniès, Benoît, et al.
Veröffentlicht: (2025) -
AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
von: Wang, Lu, et al.
Veröffentlicht: (2025) -
FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
von: Ahasan, Md Mubtasim, et al.
Veröffentlicht: (2025) -
Neural Codec Source Tracing: Toward Comprehensive Attribution in Open-Set Condition
von: Xie, Yuankun, et al.
Veröffentlicht: (2025)