MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Haina, Zhou, Yizhi, Chen, Hangting, Yu, Jianwei, Ma, Ziyang, Gu, Rongzhi, Luo, Yi, Tan, Wei, Chen, Xie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Layer-wise Investigation of Large-Scale Self-Supervised Music Representation Models
von: Zhou, Yizhi, et al.
Veröffentlicht: (2025)
von: Zhou, Yizhi, et al.
Veröffentlicht: (2025)
MuCodec: Ultra Low-Bitrate Music Codec
von: Xu, Yaoxun, et al.
Veröffentlicht: (2024)
von: Xu, Yaoxun, et al.
Veröffentlicht: (2024)
SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
WAKE: Watermarking Audio with Key Enrichment
von: Xu, Yaoxun, et al.
Veröffentlicht: (2025)
von: Xu, Yaoxun, et al.
Veröffentlicht: (2025)
AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook
von: Chen, Yushen, et al.
Veröffentlicht: (2025)
von: Chen, Yushen, et al.
Veröffentlicht: (2025)
Gull: A Generative Multifunctional Audio Codec
von: Luo, Yi, et al.
Veröffentlicht: (2024)
von: Luo, Yi, et al.
Veröffentlicht: (2024)
NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
von: Niu, Zhikang, et al.
Veröffentlicht: (2024)
von: Niu, Zhikang, et al.
Veröffentlicht: (2024)
SongPrep: A Preprocessing Framework and End-to-end Model for Full-song Structure Parsing and Lyrics Transcription
von: Tan, Wei, et al.
Veröffentlicht: (2025)
von: Tan, Wei, et al.
Veröffentlicht: (2025)
Continuous Target Speech Extraction: Enhancing Personalized Diarization and Extraction on Complex Recordings
von: Zhao, He, et al.
Veröffentlicht: (2024)
von: Zhao, He, et al.
Veröffentlicht: (2024)
ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized Transformers
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024)
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024)
The SJTU X-LANCE Lab System for MSR Challenge 2025
von: Zhu, Jinxuan, et al.
Veröffentlicht: (2026)
von: Zhu, Jinxuan, et al.
Veröffentlicht: (2026)
Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning
von: Wilkins, Julia, et al.
Veröffentlicht: (2025)
von: Wilkins, Julia, et al.
Veröffentlicht: (2025)
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
von: Wilkins, Julia, et al.
Veröffentlicht: (2024)
von: Wilkins, Julia, et al.
Veröffentlicht: (2024)
MuPT: A Generative Symbolic Music Pretrained Transformer
von: Qu, Xingwei, et al.
Veröffentlicht: (2024)
von: Qu, Xingwei, et al.
Veröffentlicht: (2024)
Progressive Residual Extraction based Pre-training for Speech Representation Learning
von: Wang, Tianrui, et al.
Veröffentlicht: (2024)
von: Wang, Tianrui, et al.
Veröffentlicht: (2024)
Mel-RoFormer for Vocal Separation and Vocal Melody Transcription
von: Wang, Ju-Chiang, et al.
Veröffentlicht: (2024)
von: Wang, Ju-Chiang, et al.
Veröffentlicht: (2024)
FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter
von: Lv, Yuanjun, et al.
Veröffentlicht: (2024)
von: Lv, Yuanjun, et al.
Veröffentlicht: (2024)
DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
von: Chen, Yafeng, et al.
Veröffentlicht: (2024)
von: Chen, Yafeng, et al.
Veröffentlicht: (2024)
Music Genre Classification: A Comparative Analysis of CNN and XGBoost Approaches with Mel-frequency cepstral coefficients and Mel Spectrograms
von: Meng, Yigang
Veröffentlicht: (2024)
von: Meng, Yigang
Veröffentlicht: (2024)
Variable Bitrate Residual Vector Quantization for Audio Coding
von: Chae, Yunkee, et al.
Veröffentlicht: (2024)
von: Chae, Yunkee, et al.
Veröffentlicht: (2024)
MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models
von: Liu, Shansong, et al.
Veröffentlicht: (2024)
von: Liu, Shansong, et al.
Veröffentlicht: (2024)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
von: Chen, Yafeng, et al.
Veröffentlicht: (2023)
von: Chen, Yafeng, et al.
Veröffentlicht: (2023)
LeVo: High-Quality Song Generation with Multi-Preference Alignment
von: Lei, Shun, et al.
Veröffentlicht: (2025)
von: Lei, Shun, et al.
Veröffentlicht: (2025)
Refining Self-Supervised Learnt Speech Representation using Brain Activations
von: Li, Hengyu, et al.
Veröffentlicht: (2024)
von: Li, Hengyu, et al.
Veröffentlicht: (2024)
AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
Pianoroll-Event: A Novel Score Representation for Symbolic Music
von: Qian, Lekai, et al.
Veröffentlicht: (2026)
von: Qian, Lekai, et al.
Veröffentlicht: (2026)
Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
von: Guo, Hongming, et al.
Veröffentlicht: (2024)
von: Guo, Hongming, et al.
Veröffentlicht: (2024)
A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations
von: Chiu, Aemon Yat Fei, et al.
Veröffentlicht: (2025)
von: Chiu, Aemon Yat Fei, et al.
Veröffentlicht: (2025)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
MoMu-Diffusion: On Learning Long-Term Motion-Music Synchronization and Correspondence
von: You, Fuming, et al.
Veröffentlicht: (2024)
von: You, Fuming, et al.
Veröffentlicht: (2024)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
Mel-Spectrogram Inversion via Alternating Direction Method of Multipliers
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
MelTok: 2D Tokenization for Single-Codebook Audio Compression
von: Li, Jingyi, et al.
Veröffentlicht: (2025)
von: Li, Jingyi, et al.
Veröffentlicht: (2025)
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
Vector Quantized Diffusion Model Based Speech Bandwidth Extension
von: Fang, Yuan, et al.
Veröffentlicht: (2024)
von: Fang, Yuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Layer-wise Investigation of Large-Scale Self-Supervised Music Representation Models
von: Zhou, Yizhi, et al.
Veröffentlicht: (2025) -
MuCodec: Ultra Low-Bitrate Music Codec
von: Xu, Yaoxun, et al.
Veröffentlicht: (2024) -
SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor
von: Yang, Chenyu, et al.
Veröffentlicht: (2024) -
WAKE: Watermarking Audio with Key Enrichment
von: Xu, Yaoxun, et al.
Veröffentlicht: (2025) -
AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook
von: Chen, Yushen, et al.
Veröffentlicht: (2025)