Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Wilkins, Julia, Ding, Sivan, Fuentes, Magdalena, Bello, Juan Pablo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
por: Wilkins, Julia, et al.
Publicado: (2024)
por: Wilkins, Julia, et al.
Publicado: (2024)
Do Music Source Separation Models Preserve Spatial Information in Binaural Audio?
por: Namballa, Richa, et al.
Publicado: (2025)
por: Namballa, Richa, et al.
Publicado: (2025)
Learning Multidimensional Disentangled Representations of Instrumental Sounds for Musical Similarity Assessment
por: Hashizume, Yuka, et al.
Publicado: (2024)
por: Hashizume, Yuka, et al.
Publicado: (2024)
Bird Vocalization Embedding Extraction Using Self-Supervised Disentangled Representation Learning
por: Shi, Runwu, et al.
Publicado: (2024)
por: Shi, Runwu, et al.
Publicado: (2024)
SONIQUE: Video Background Music Generation Using Unpaired Audio-Visual Data
por: Zhang, Liqian, et al.
Publicado: (2024)
por: Zhang, Liqian, et al.
Publicado: (2024)
Musical Source Separation of Brazilian Percussion
por: Namballa, Richa, et al.
Publicado: (2025)
por: Namballa, Richa, et al.
Publicado: (2025)
Quantifying Dimensional Independence in Speech: An Information-Theoretic Framework for Disentangled Representation Learning
por: Kashyap, Bipasha, et al.
Publicado: (2026)
por: Kashyap, Bipasha, et al.
Publicado: (2026)
Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
por: Liu, Rui, et al.
Publicado: (2024)
por: Liu, Rui, et al.
Publicado: (2024)
Evaluating Disentangled Representations for Controllable Music Generation
por: Ibáñez-Martínez, Laura, et al.
Publicado: (2026)
por: Ibáñez-Martínez, Laura, et al.
Publicado: (2026)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
por: Chen, Yafeng, et al.
Publicado: (2024)
por: Chen, Yafeng, et al.
Publicado: (2024)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
por: Chen, Yafeng, et al.
Publicado: (2023)
por: Chen, Yafeng, et al.
Publicado: (2023)
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
por: Nam, KiHyun, et al.
Publicado: (2024)
por: Nam, KiHyun, et al.
Publicado: (2024)
Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant Retrieval
por: Deng, Yimin, et al.
Publicado: (2024)
por: Deng, Yimin, et al.
Publicado: (2024)
MEDIC: Zero-shot Music Editing with Disentangled Inversion Control
por: Liu, Huadai, et al.
Publicado: (2024)
por: Liu, Huadai, et al.
Publicado: (2024)
Learning Separated Representations for Instrument-based Music Similarity
por: Hashizume, Yuka, et al.
Publicado: (2025)
por: Hashizume, Yuka, et al.
Publicado: (2025)
VISinger2+: End-to-End Singing Voice Synthesis Augmented by Self-Supervised Learning Representation
por: Yu, Yifeng, et al.
Publicado: (2024)
por: Yu, Yifeng, et al.
Publicado: (2024)
Layer-wise Investigation of Large-Scale Self-Supervised Music Representation Models
por: Zhou, Yizhi, et al.
Publicado: (2025)
por: Zhou, Yizhi, et al.
Publicado: (2025)
Post-Training Quantization for Audio Diffusion Transformers
por: Khandelwal, Tanmay, et al.
Publicado: (2025)
por: Khandelwal, Tanmay, et al.
Publicado: (2025)
Self-Supervised Learning of Spatial Acoustic Representation with Cross-Channel Signal Reconstruction and Multi-Channel Conformer
por: Yang, Bing, et al.
Publicado: (2023)
por: Yang, Bing, et al.
Publicado: (2023)
Refining Self-Supervised Learnt Speech Representation using Brain Activations
por: Li, Hengyu, et al.
Publicado: (2024)
por: Li, Hengyu, et al.
Publicado: (2024)
HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement
por: Hussein, Amir, et al.
Publicado: (2025)
por: Hussein, Amir, et al.
Publicado: (2025)
Semi-Supervised Self-Learning Enhanced Music Emotion Recognition
por: Sun, Yifu, et al.
Publicado: (2024)
por: Sun, Yifu, et al.
Publicado: (2024)
Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment
por: Stahl, Benjamin, et al.
Publicado: (2025)
por: Stahl, Benjamin, et al.
Publicado: (2025)
Adapting General Disentanglement-Based Speaker Anonymization for Enhanced Emotion Preservation
por: Miao, Xiaoxiao, et al.
Publicado: (2024)
por: Miao, Xiaoxiao, et al.
Publicado: (2024)
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
por: Deng, Yimin, et al.
Publicado: (2024)
por: Deng, Yimin, et al.
Publicado: (2024)
Leveraging Self-Supervised Learning for Speaker Diarization
por: Han, Jiangyu, et al.
Publicado: (2024)
por: Han, Jiangyu, et al.
Publicado: (2024)
Self-Supervised Disentangled Representation Learning for Robust Target Speech Extraction
por: Mu, Zhaoxi, et al.
Publicado: (2023)
por: Mu, Zhaoxi, et al.
Publicado: (2023)
OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction
por: Alonso-Jiménez, Pablo, et al.
Publicado: (2025)
por: Alonso-Jiménez, Pablo, et al.
Publicado: (2025)
Music Similarity Representation Learning Focusing on Individual Instruments with Source Separation and Human Preference
por: Imamura, Takehiro, et al.
Publicado: (2025)
por: Imamura, Takehiro, et al.
Publicado: (2025)
Latent Acoustic Mapping for Direction of Arrival Estimation: A Self-Supervised Approach
por: Roman, Adrian S., et al.
Publicado: (2025)
por: Roman, Adrian S., et al.
Publicado: (2025)
A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations
por: Chiu, Aemon Yat Fei, et al.
Publicado: (2025)
por: Chiu, Aemon Yat Fei, et al.
Publicado: (2025)
Speaker Recognition Using Isomorphic Graph Attention Network Based Pooling on Self-Supervised Representation
por: Ge, Zirui, et al.
Publicado: (2023)
por: Ge, Zirui, et al.
Publicado: (2023)
Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval
por: Stewart, Shanti, et al.
Publicado: (2024)
por: Stewart, Shanti, et al.
Publicado: (2024)
Learning Disentangled Speech Representations
por: Brima, Yusuf, et al.
Publicado: (2023)
por: Brima, Yusuf, et al.
Publicado: (2023)
Low-Resource Self-Supervised Learning with SSL-Enhanced TTS
por: Hsu, Po-chun, et al.
Publicado: (2023)
por: Hsu, Po-chun, et al.
Publicado: (2023)
Geometric Analysis of Speech Representation Spaces: Topological Disentanglement and Confound Detection
por: Kashyap, Bipasha, et al.
Publicado: (2026)
por: Kashyap, Bipasha, et al.
Publicado: (2026)
Emotion-driven Piano Music Generation via Two-stage Disentanglement and Functional Representation
por: Huang, Jingyue, et al.
Publicado: (2024)
por: Huang, Jingyue, et al.
Publicado: (2024)
Pianoroll-Event: A Novel Score Representation for Symbolic Music
por: Qian, Lekai, et al.
Publicado: (2026)
por: Qian, Lekai, et al.
Publicado: (2026)
SLAP: Learning Speaker and Health-Related Representations from Natural Language Supervision
por: Ando, Angelika, et al.
Publicado: (2025)
por: Ando, Angelika, et al.
Publicado: (2025)
MusicAOG: an Energy-Based Model for Learning and Sampling a Hierarchical Representation of Symbolic Music
por: Qian, Yikai, et al.
Publicado: (2024)
por: Qian, Yikai, et al.
Publicado: (2024)
Ejemplares similares
-
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
por: Wilkins, Julia, et al.
Publicado: (2024) -
Do Music Source Separation Models Preserve Spatial Information in Binaural Audio?
por: Namballa, Richa, et al.
Publicado: (2025) -
Learning Multidimensional Disentangled Representations of Instrumental Sounds for Musical Similarity Assessment
por: Hashizume, Yuka, et al.
Publicado: (2024) -
Bird Vocalization Embedding Extraction Using Self-Supervised Disentangled Representation Learning
por: Shi, Runwu, et al.
Publicado: (2024) -
SONIQUE: Video Background Music Generation Using Unpaired Audio-Visual Data
por: Zhang, Liqian, et al.
Publicado: (2024)