Variable Bitrate Residual Vector Quantization for Audio Coding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chae, Yunkee, Choi, Woosung, Takida, Yuhta, Koo, Junghyun, Ikemiya, Yukara, Zhong, Zhi, Cheuk, Kin Wai, Martínez-Ramírez, Marco A., Lee, Kyogu, Liao, Wei-Hsiang, Mitsufuji, Yuki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Large-Scale Training Data Attribution for Music Generative Models via Unlearning
von: Choi, Woosung, et al.
Veröffentlicht: (2025)
von: Choi, Woosung, et al.
Veröffentlicht: (2025)
Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
von: Chae, Yunkee, et al.
Veröffentlicht: (2025)
von: Chae, Yunkee, et al.
Veröffentlicht: (2025)
Concept-TRAK: Understanding how diffusion models learn concepts through concept-level attribution
von: Park, Yonghyun, et al.
Veröffentlicht: (2025)
von: Park, Yonghyun, et al.
Veröffentlicht: (2025)
Break-the-Beat! Controllable MIDI-to-Drum Audio Synthesis
von: Cui, Shuyang, et al.
Veröffentlicht: (2026)
von: Cui, Shuyang, et al.
Veröffentlicht: (2026)
HQ-VAE: Hierarchical Discrete Representation Learning with Variational Bayes
von: Takida, Yuhta, et al.
Veröffentlicht: (2023)
von: Takida, Yuhta, et al.
Veröffentlicht: (2023)
MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source Extraction
von: Chae, Yunkee, et al.
Veröffentlicht: (2025)
von: Chae, Yunkee, et al.
Veröffentlicht: (2025)
CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation
von: Chen, Yuanhong, et al.
Veröffentlicht: (2025)
von: Chen, Yuanhong, et al.
Veröffentlicht: (2025)
DisMix: Disentangling Mixtures of Musical Instruments for Source-level Pitch and Timbre Manipulation
von: Luo, Yin-Jyun, et al.
Veröffentlicht: (2024)
von: Luo, Yin-Jyun, et al.
Veröffentlicht: (2024)
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023)
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023)
Music Foundation Model as Generic Booster for Music Downstream Tasks
von: Liao, WeiHsiang, et al.
Veröffentlicht: (2024)
von: Liao, WeiHsiang, et al.
Veröffentlicht: (2024)
Timbre-Trap: A Low-Resource Framework for Instrument-Agnostic Music Transcription
von: Cwitkowitz, Frank, et al.
Veröffentlicht: (2023)
von: Cwitkowitz, Frank, et al.
Veröffentlicht: (2023)
Improving Vector-Quantized Image Modeling with Latent Consistency-Matching Diffusion
von: Nguyen, Bac, et al.
Veröffentlicht: (2024)
von: Nguyen, Bac, et al.
Veröffentlicht: (2024)
Efficiency without Compromise: CLIP-aided Text-to-Image GANs with Increased Diversity
von: Kobayashi, Yuya, et al.
Veröffentlicht: (2025)
von: Kobayashi, Yuya, et al.
Veröffentlicht: (2025)
Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tuning
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
Automatic Music Mixing using a Generative Model of Effect Embeddings
von: Moliner, Eloi, et al.
Veröffentlicht: (2025)
von: Moliner, Eloi, et al.
Veröffentlicht: (2025)
PAVAS: Physics-Aware Video-to-Audio Synthesis
von: Hyun-Bin, Oh, et al.
Veröffentlicht: (2025)
von: Hyun-Bin, Oh, et al.
Veröffentlicht: (2025)
Towards Blind Data Cleaning: A Case Study in Music Source Separation
von: Gui, Azalea, et al.
Veröffentlicht: (2025)
von: Gui, Azalea, et al.
Veröffentlicht: (2025)
Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer
von: Mancusi, Michele, et al.
Veröffentlicht: (2024)
von: Mancusi, Michele, et al.
Veröffentlicht: (2024)
Denoising Multi-Beta VAE: Representation Learning for Disentanglement and Generation
von: Uppal, Anshuk, et al.
Veröffentlicht: (2025)
von: Uppal, Anshuk, et al.
Veröffentlicht: (2025)
MR-MT3: Memory Retaining Multi-Track Music Transcription to Mitigate Instrument Leakage
von: Tan, Hao Hao, et al.
Veröffentlicht: (2024)
von: Tan, Hao Hao, et al.
Veröffentlicht: (2024)
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
von: Comunità, Marco, et al.
Veröffentlicht: (2024)
von: Comunità, Marco, et al.
Veröffentlicht: (2024)
LLM2Fx-Tools: Tool Calling For Music Post-Production
von: Doh, Seungheon, et al.
Veröffentlicht: (2025)
von: Doh, Seungheon, et al.
Veröffentlicht: (2025)
A Comprehensive Real-World Assessment of Audio Watermarking Algorithms: Will They Survive Neural Codecs?
von: Özer, Yigitcan, et al.
Veröffentlicht: (2025)
von: Özer, Yigitcan, et al.
Veröffentlicht: (2025)
ITO-Master: Inference-Time Optimization for Audio Effects Modeling of Music Mastering Processors
von: Koo, Junghyun, et al.
Veröffentlicht: (2025)
von: Koo, Junghyun, et al.
Veröffentlicht: (2025)
Can Large Language Models Predict Audio Effects Parameters from Natural Language?
von: Doh, Seungheon, et al.
Veröffentlicht: (2025)
von: Doh, Seungheon, et al.
Veröffentlicht: (2025)
Transformed Low-rank Adaptation via Tensor Decomposition and Its Applications to Text-to-image Models
von: Tao, Zerui, et al.
Veröffentlicht: (2025)
von: Tao, Zerui, et al.
Veröffentlicht: (2025)
Demystifying MaskGIT Sampler and Beyond: Adaptive Order Selection in Masked Diffusion
von: Hayakawa, Satoshi, et al.
Veröffentlicht: (2025)
von: Hayakawa, Satoshi, et al.
Veröffentlicht: (2025)
Distillation of Discrete Diffusion through Dimensional Correlations
von: Hayakawa, Satoshi, et al.
Veröffentlicht: (2024)
von: Hayakawa, Satoshi, et al.
Veröffentlicht: (2024)
MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
Improving Unsupervised Clean-to-Rendered Guitar Tone Transformation Using GANs and Integrated Unaligned Clean Data
von: Chen, Yu-Hua, et al.
Veröffentlicht: (2024)
von: Chen, Yu-Hua, et al.
Veröffentlicht: (2024)
Distill, Forget, Repeat: A Framework for Continual Unlearning in Text-to-Image Diffusion Models
von: George, Naveen, et al.
Veröffentlicht: (2025)
von: George, Naveen, et al.
Veröffentlicht: (2025)
VCT: Training Consistency Models with Variational Noise Coupling
von: Silvestri, Gianluigi, et al.
Veröffentlicht: (2025)
von: Silvestri, Gianluigi, et al.
Veröffentlicht: (2025)
Fx-Encoder++: Extracting Instrument-Wise Audio Effects Representations from Mixtures
von: Yeh, Yen-Tung, et al.
Veröffentlicht: (2025)
von: Yeh, Yen-Tung, et al.
Veröffentlicht: (2025)
DDD: A Perceptually Superior Low-Response-Time DNN-based Declipper
von: Yi, Jayeon, et al.
Veröffentlicht: (2024)
von: Yi, Jayeon, et al.
Veröffentlicht: (2024)
Song Form-aware Full-Song Text-to-Lyrics Generation with Multi-Level Granularity Syllable Count Control
von: Chae, Yunkee, et al.
Veröffentlicht: (2024)
von: Chae, Yunkee, et al.
Veröffentlicht: (2024)
Theoretical Refinement of CLIP by Utilizing Linear Structure of Optimal Similarity
von: Yoshida, Naoki, et al.
Veröffentlicht: (2025)
von: Yoshida, Naoki, et al.
Veröffentlicht: (2025)
$\textit{Jump Your Steps}$: Optimizing Sampling Schedule of Discrete Diffusion Models
von: Park, Yong-Hyun, et al.
Veröffentlicht: (2024)
von: Park, Yong-Hyun, et al.
Veröffentlicht: (2024)
PaGoDA: Progressive Growing of a One-Step Generator from a Low-Resolution Diffusion Teacher
von: Kim, Dongjun, et al.
Veröffentlicht: (2024)
von: Kim, Dongjun, et al.
Veröffentlicht: (2024)
GRAFX: An Open-Source Library for Audio Processing Graphs in PyTorch
von: Lee, Sungho, et al.
Veröffentlicht: (2024)
von: Lee, Sungho, et al.
Veröffentlicht: (2024)
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Large-Scale Training Data Attribution for Music Generative Models via Unlearning
von: Choi, Woosung, et al.
Veröffentlicht: (2025) -
Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
von: Chae, Yunkee, et al.
Veröffentlicht: (2025) -
Concept-TRAK: Understanding how diffusion models learn concepts through concept-level attribution
von: Park, Yonghyun, et al.
Veröffentlicht: (2025) -
Break-the-Beat! Controllable MIDI-to-Drum Audio Synthesis
von: Cui, Shuyang, et al.
Veröffentlicht: (2026) -
HQ-VAE: Hierarchical Discrete Representation Learning with Variational Bayes
von: Takida, Yuhta, et al.
Veröffentlicht: (2023)