Learning Source Disentanglement in Neural Audio Codec
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bie, Xiaoyu, Liu, Xubo, Richard, Gaël |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
Désentrelacement Fréquentiel Doux pour les Codecs Audio Neuronaux
von: Giniès, Benoît, et al.
Veröffentlicht: (2025)
von: Giniès, Benoît, et al.
Veröffentlicht: (2025)
CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio
von: Pasini, Marco, et al.
Veröffentlicht: (2025)
von: Pasini, Marco, et al.
Veröffentlicht: (2025)
Gull: A Generative Multifunctional Audio Codec
von: Luo, Yi, et al.
Veröffentlicht: (2024)
von: Luo, Yi, et al.
Veröffentlicht: (2024)
Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
PoDAR: Power-Disentangled Audio Representation for Generative Modeling
von: Luebs, Alejandro, et al.
Veröffentlicht: (2026)
von: Luebs, Alejandro, et al.
Veröffentlicht: (2026)
SNAC: Multi-Scale Neural Audio Codec
von: Siuzdak, Hubert, et al.
Veröffentlicht: (2024)
von: Siuzdak, Hubert, et al.
Veröffentlicht: (2024)
Translation-Equivariant Self-Supervised Learning for Pitch Estimation with Optimal Transport
von: Torres, Bernardo, et al.
Veröffentlicht: (2025)
von: Torres, Bernardo, et al.
Veröffentlicht: (2025)
MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation
von: Song, Yakun, et al.
Veröffentlicht: (2025)
von: Song, Yakun, et al.
Veröffentlicht: (2025)
A Comprehensive Real-World Assessment of Audio Watermarking Algorithms: Will They Survive Neural Codecs?
von: Özer, Yigitcan, et al.
Veröffentlicht: (2025)
von: Özer, Yigitcan, et al.
Veröffentlicht: (2025)
SpectroStream: A Versatile Neural Codec for General Audio
von: Li, Yunpeng, et al.
Veröffentlicht: (2025)
von: Li, Yunpeng, et al.
Veröffentlicht: (2025)
F-StrIPE: Fast Structure-Informed Positional Encoding for Symbolic Music Generation
von: Agarwal, Manvi, et al.
Veröffentlicht: (2025)
von: Agarwal, Manvi, et al.
Veröffentlicht: (2025)
TACNET: Temporal Audio Source Counting Network
von: Ahmadnejad, Amirreza, et al.
Veröffentlicht: (2023)
von: Ahmadnejad, Amirreza, et al.
Veröffentlicht: (2023)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
Training-Free Multi-Step Audio Source Separation
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
On Temporal Guidance and Iterative Refinement in Audio Source Separation
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
Text-Queried Audio Source Separation via Hierarchical Modeling
von: Yin, Xinlei, et al.
Veröffentlicht: (2025)
von: Yin, Xinlei, et al.
Veröffentlicht: (2025)
Masked Audio Generation using a Single Non-Autoregressive Transformer
von: Ziv, Alon, et al.
Veröffentlicht: (2024)
von: Ziv, Alon, et al.
Veröffentlicht: (2024)
DisMix: Disentangling Mixtures of Musical Instruments for Source-level Pitch and Timbre Manipulation
von: Luo, Yin-Jyun, et al.
Veröffentlicht: (2024)
von: Luo, Yin-Jyun, et al.
Veröffentlicht: (2024)
AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs
von: Nguyen, Hoang, et al.
Veröffentlicht: (2025)
von: Nguyen, Hoang, et al.
Veröffentlicht: (2025)
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
von: Wang, Yuancheng, et al.
Veröffentlicht: (2024)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2024)
Facing the Music: Tackling Singing Voice Separation in Cinematic Audio Source Separation
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2024)
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2024)
FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
Neural Codec Source Tracing: Toward Comprehensive Attribution in Open-Set Condition
von: Xie, Yuankun, et al.
Veröffentlicht: (2025)
von: Xie, Yuankun, et al.
Veröffentlicht: (2025)
Quantifying Multimodal Imbalance: A GMM-Guided Adaptive Loss for Audio-Visual Learning
von: Liu, Zhaocheng, et al.
Veröffentlicht: (2025)
von: Liu, Zhaocheng, et al.
Veröffentlicht: (2025)
FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
Towards Neural Audio Codec Source Parsing
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Latent Granular Resynthesis using Neural Audio Codecs
von: Tokui, Nao, et al.
Veröffentlicht: (2025)
von: Tokui, Nao, et al.
Veröffentlicht: (2025)
Towards Audio Codec-based Speech Separation
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
Self-Supervised Disentangled Representation Learning for Robust Target Speech Extraction
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2023)
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2023)
Generating Sample-Based Musical Instruments Using Neural Audio Codec Language Models
von: Nercessian, Shahan, et al.
Veröffentlicht: (2024)
von: Nercessian, Shahan, et al.
Veröffentlicht: (2024)
AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
A Survey of Deep Learning Audio Generation Methods
von: Božić, Matej, et al.
Veröffentlicht: (2024)
von: Božić, Matej, et al.
Veröffentlicht: (2024)
DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
Guiding Audio Editing with Audio Language Model
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
Codec-Based Deepfake Source Tracing via Neural Audio Codec Taxonomy
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
von: Li, Haoyang, et al.
Veröffentlicht: (2025) -
Désentrelacement Fréquentiel Doux pour les Codecs Audio Neuronaux
von: Giniès, Benoît, et al.
Veröffentlicht: (2025) -
CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio
von: Pasini, Marco, et al.
Veröffentlicht: (2025) -
Gull: A Generative Multifunctional Audio Codec
von: Luo, Yi, et al.
Veröffentlicht: (2024) -
Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)