Linear Complexity Self-Supervised Learning for Music Understanding with Random Quantizer
Fuente:
arXiv
Guardado en:
| Autores principales: | Vavaroutsos, Petros, Palamas, Theodoros, Vikatos, Pantelis |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Latent Space Disentanglement via Activation Steering for Interpretable Attribute Control in Symbolic Music Generation
por: Prokopiou, Ioannis, et al.
Publicado: (2026)
por: Prokopiou, Ioannis, et al.
Publicado: (2026)
LabelBuddy: An Open Source Music and Audio Language Annotation Tagging Tool Using AI Assistance
por: Prokopiou, Ioannis, et al.
Publicado: (2026)
por: Prokopiou, Ioannis, et al.
Publicado: (2026)
MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
por: Zhu, Haina, et al.
Publicado: (2025)
por: Zhu, Haina, et al.
Publicado: (2025)
MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training
por: Li, Yizhi, et al.
Publicado: (2023)
por: Li, Yizhi, et al.
Publicado: (2023)
Music for All: Representational Bias and Cross-Cultural Adaptability of Music Generation Models
por: Mehta, Atharva, et al.
Publicado: (2025)
por: Mehta, Atharva, et al.
Publicado: (2025)
MusicLIME: Explainable Multimodal Music Understanding
por: Sotirou, Theodoros, et al.
Publicado: (2024)
por: Sotirou, Theodoros, et al.
Publicado: (2024)
ChatMusician: Understanding and Generating Music Intrinsically with LLM
por: Yuan, Ruibin, et al.
Publicado: (2024)
por: Yuan, Ruibin, et al.
Publicado: (2024)
AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?
por: Ok, Hyunjong, et al.
Publicado: (2025)
por: Ok, Hyunjong, et al.
Publicado: (2025)
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
por: Seth, Ashish, et al.
Publicado: (2024)
por: Seth, Ashish, et al.
Publicado: (2024)
EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
por: Chen, Wenxi, et al.
Publicado: (2024)
por: Chen, Wenxi, et al.
Publicado: (2024)
Do Music Generation Models Encode Music Theory?
por: Wei, Megan, et al.
Publicado: (2024)
por: Wei, Megan, et al.
Publicado: (2024)
CSyMR: Benchmarking Compositional Music Information Retrieval in Symbolic Music Reasoning
por: Wang, Boyang, et al.
Publicado: (2025)
por: Wang, Boyang, et al.
Publicado: (2025)
Joint Fine-tuning and Conversion of Pretrained Speech and Language Models towards Linear Complexity
por: He, Mutian, et al.
Publicado: (2024)
por: He, Mutian, et al.
Publicado: (2024)
Learning When to Think While Listening in Large Audio-Language Models
por: Song, Zhiyuan, et al.
Publicado: (2026)
por: Song, Zhiyuan, et al.
Publicado: (2026)
Foundation Models for Music: A Survey
por: Ma, Yinghao, et al.
Publicado: (2024)
por: Ma, Yinghao, et al.
Publicado: (2024)
Large Language Models' Internal Perception of Symbolic Music
por: Shin, Andrew, et al.
Publicado: (2025)
por: Shin, Andrew, et al.
Publicado: (2025)
Missing Melodies: AI Music Generation and its "Nearly" Complete Omission of the Global South
por: Mehta, Atharva, et al.
Publicado: (2024)
por: Mehta, Atharva, et al.
Publicado: (2024)
Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation
por: Lokegaonkar, Vaibhavi, et al.
Publicado: (2026)
por: Lokegaonkar, Vaibhavi, et al.
Publicado: (2026)
MAEB: Massive Audio Embedding Benchmark
por: Assadi, Adnan El, et al.
Publicado: (2026)
por: Assadi, Adnan El, et al.
Publicado: (2026)
DiffuSpeech: Silent Thought, Spoken Answer via Unified Speech-Text Diffusion
por: Lou, Yuxuan, et al.
Publicado: (2026)
por: Lou, Yuxuan, et al.
Publicado: (2026)
MoST: Mixing Speech and Text with Modality-Aware Mixture of Experts
por: Lou, Yuxuan, et al.
Publicado: (2026)
por: Lou, Yuxuan, et al.
Publicado: (2026)
WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation
por: Della Libera, Luca, et al.
Publicado: (2026)
por: Della Libera, Luca, et al.
Publicado: (2026)
HyperCLOVA X 8B Omni
por: NAVER Cloud HyperCLOVA X Team
Publicado: (2026)
por: NAVER Cloud HyperCLOVA X Team
Publicado: (2026)
EchoChain: A Full-Duplex Benchmark for State-Update Reasoning Under Interruptions
por: Modi, Smit Nautambhai, et al.
Publicado: (2026)
por: Modi, Smit Nautambhai, et al.
Publicado: (2026)
Speech Emotion Recognition Leveraging OpenAI's Whisper Representations and Attentive Pooling Methods
por: Shendabadi, Ali, et al.
Publicado: (2026)
por: Shendabadi, Ali, et al.
Publicado: (2026)
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
por: Bogavelli, Tara, et al.
Publicado: (2026)
por: Bogavelli, Tara, et al.
Publicado: (2026)
Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data
por: Kumar, Gokul Karthik, et al.
Publicado: (2025)
por: Kumar, Gokul Karthik, et al.
Publicado: (2025)
PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation
por: He, Jiajun, et al.
Publicado: (2025)
por: He, Jiajun, et al.
Publicado: (2025)
Synthetic Audio Helps for Cognitive State Tasks
por: Soubki, Adil, et al.
Publicado: (2025)
por: Soubki, Adil, et al.
Publicado: (2025)
Exploring Adapter Design Tradeoffs for Low Resource Music Generation
por: Mehta, Atharva, et al.
Publicado: (2025)
por: Mehta, Atharva, et al.
Publicado: (2025)
ComposerX: Multi-Agent Symbolic Music Composition with LLMs
por: Deng, Qixin, et al.
Publicado: (2024)
por: Deng, Qixin, et al.
Publicado: (2024)
QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding
por: Biswas, Subrata, et al.
Publicado: (2025)
por: Biswas, Subrata, et al.
Publicado: (2025)
Advancing Speech Understanding in Speech-Aware Language Models with GRPO
por: Elmakies, Avishai, et al.
Publicado: (2025)
por: Elmakies, Avishai, et al.
Publicado: (2025)
Myna: Masking-Based Contrastive Learning of Musical Representations
por: Yonay, Ori, et al.
Publicado: (2025)
por: Yonay, Ori, et al.
Publicado: (2025)
Automatic Time Signature Determination for New Scores Using Lyrics for Latent Rhythmic Structure
por: Liao, Callie C., et al.
Publicado: (2023)
por: Liao, Callie C., et al.
Publicado: (2023)
LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model
por: Sun, Yirong, et al.
Publicado: (2025)
por: Sun, Yirong, et al.
Publicado: (2025)
Melody or Machine: Detecting Synthetic Music with Dual-Stream Contrastive Learning
por: Batra, Arnesh, et al.
Publicado: (2025)
por: Batra, Arnesh, et al.
Publicado: (2025)
Improving Self-supervised Pre-training using Accent-Specific Codebooks
por: Prabhu, Darshan, et al.
Publicado: (2024)
por: Prabhu, Darshan, et al.
Publicado: (2024)
Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation Models
por: Hu, Yuchen, et al.
Publicado: (2024)
por: Hu, Yuchen, et al.
Publicado: (2024)
wav2graph: A Framework for Supervised Learning Knowledge Graph from Speech
por: Le-Duc, Khai, et al.
Publicado: (2024)
por: Le-Duc, Khai, et al.
Publicado: (2024)
Ejemplares similares
-
Latent Space Disentanglement via Activation Steering for Interpretable Attribute Control in Symbolic Music Generation
por: Prokopiou, Ioannis, et al.
Publicado: (2026) -
LabelBuddy: An Open Source Music and Audio Language Annotation Tagging Tool Using AI Assistance
por: Prokopiou, Ioannis, et al.
Publicado: (2026) -
MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
por: Zhu, Haina, et al.
Publicado: (2025) -
MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training
por: Li, Yizhi, et al.
Publicado: (2023) -
Music for All: Representational Bias and Cross-Cultural Adaptability of Music Generation Models
por: Mehta, Atharva, et al.
Publicado: (2025)