Evaluating Multimodal Large Language Models on Core Music Perception Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Carone, Brandon James, Roman, Iran R., Ripollés, Pablo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS
by: Carone, Brandon James, et al.
Published: (2025)
by: Carone, Brandon James, et al.
Published: (2025)
SoundSignature: What Type of Music Do You Like?
by: Carone, Brandon James, et al.
Published: (2024)
by: Carone, Brandon James, et al.
Published: (2024)
The Music Maestro or The Musically Challenged, A Massive Music Evaluation Benchmark for Large Language Models
by: Li, Jiajia, et al.
Published: (2024)
by: Li, Jiajia, et al.
Published: (2024)
Latent Acoustic Mapping for Direction of Arrival Estimation: A Self-Supervised Approach
by: Roman, Adrian S., et al.
Published: (2025)
by: Roman, Adrian S., et al.
Published: (2025)
Content-based Controls For Music Large Language Modeling
by: Lin, Liwei, et al.
Published: (2023)
by: Lin, Liwei, et al.
Published: (2023)
NotaGen: Advancing Musicality in Symbolic Music Generation with Large Language Model Training Paradigms
by: Wang, Yashan, et al.
Published: (2025)
by: Wang, Yashan, et al.
Published: (2025)
Language Model Mapping in Multimodal Music Learning: A Grand Challenge Proposal
by: Chin, Daniel, et al.
Published: (2025)
by: Chin, Daniel, et al.
Published: (2025)
M2M-Gen: A Multimodal Framework for Automated Background Music Generation in Japanese Manga Using Large Language Models
by: Sharma, Megha, et al.
Published: (2024)
by: Sharma, Megha, et al.
Published: (2024)
Practical and Reproducible Symbolic Music Generation by Large Language Models with Structural Embeddings
by: Rhyu, Seungyeon, et al.
Published: (2024)
by: Rhyu, Seungyeon, et al.
Published: (2024)
Perception-Inspired Graph Convolution for Music Understanding Tasks
by: Karystinaios, Emmanouil, et al.
Published: (2024)
by: Karystinaios, Emmanouil, et al.
Published: (2024)
LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection
by: Chou, Benjamin Shiue-Hal, et al.
Published: (2025)
by: Chou, Benjamin Shiue-Hal, et al.
Published: (2025)
AHAMask: Reliable Task Specification for Large Audio Language Models without Instructions
by: Guo, Yiwei, et al.
Published: (2025)
by: Guo, Yiwei, et al.
Published: (2025)
Exploring State-Space-Model based Language Model in Music Generation
by: Lee, Wei-Jaw, et al.
Published: (2025)
by: Lee, Wei-Jaw, et al.
Published: (2025)
Video2Music: Suitable Music Generation from Videos using an Affective Multimodal Transformer model
by: Kang, Jaeyong, et al.
Published: (2023)
by: Kang, Jaeyong, et al.
Published: (2023)
Musical ethnocentrism in Large Language Models
by: Kruspe, Anna
Published: (2025)
by: Kruspe, Anna
Published: (2025)
Layer-wise Investigation of Large-Scale Self-Supervised Music Representation Models
by: Zhou, Yizhi, et al.
Published: (2025)
by: Zhou, Yizhi, et al.
Published: (2025)
Aligning Text-to-Music Evaluation with Human Preferences
by: Huang, Yichen, et al.
Published: (2025)
by: Huang, Yichen, et al.
Published: (2025)
Music Consistency Models
by: Fei, Zhengcong, et al.
Published: (2024)
by: Fei, Zhengcong, et al.
Published: (2024)
WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations
by: Kim, Jaeyeon, et al.
Published: (2025)
by: Kim, Jaeyeon, et al.
Published: (2025)
Large Language Models' Internal Perception of Symbolic Music
by: Shin, Andrew, et al.
Published: (2025)
by: Shin, Andrew, et al.
Published: (2025)
Hear: Hierarchically Enhanced Aesthetic Representations For Multidimensional Music Evaluation
by: Liu, Shuyang, et al.
Published: (2025)
by: Liu, Shuyang, et al.
Published: (2025)
VocalAgent: Large Language Models for Vocal Health Diagnostics with Safety-Aware Evaluation
by: Kim, Yubin, et al.
Published: (2025)
by: Kim, Yubin, et al.
Published: (2025)
BNMusic: Blending Environmental Noises into Personalized Music
by: Zuo, Chi, et al.
Published: (2025)
by: Zuo, Chi, et al.
Published: (2025)
Performance Improvement of Language-Queried Audio Source Separation Based on Caption Augmentation From Large Language Models for DCASE Challenge 2024 Task 9
by: Lee, Do Hyun, et al.
Published: (2024)
by: Lee, Do Hyun, et al.
Published: (2024)
AImoclips: A Benchmark for Evaluating Emotion Conveyance in Text-to-Music Generation
by: Go, Gyehun, et al.
Published: (2025)
by: Go, Gyehun, et al.
Published: (2025)
MusicFlow: Cascaded Flow Matching for Text Guided Music Generation
by: Prajwal, K R, et al.
Published: (2024)
by: Prajwal, K R, et al.
Published: (2024)
Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models
by: Yang, Wanqi, et al.
Published: (2024)
by: Yang, Wanqi, et al.
Published: (2024)
Mozart's Touch: A Lightweight Multi-modal Music Generation Framework Based on Pre-Trained Large Models
by: Li, Jiajun, et al.
Published: (2024)
by: Li, Jiajun, et al.
Published: (2024)
Guitar-TECHS: An Electric Guitar Dataset Covering Techniques, Musical Excerpts, Chords and Scales Using a Diverse Array of Hardware
by: Pedroza, Hegel, et al.
Published: (2025)
by: Pedroza, Hegel, et al.
Published: (2025)
Exploring Musical Roots: Applying Audio Embeddings to Empower Influence Attribution for a Generative Music Model
by: Barnett, Julia, et al.
Published: (2024)
by: Barnett, Julia, et al.
Published: (2024)
The Interpretation Gap in Text-to-Music Generation Models
by: Zang, Yongyi, et al.
Published: (2024)
by: Zang, Yongyi, et al.
Published: (2024)
Play Me Something Icy: Practical Challenges, Explainability and the Semantic Gap in Generative AI Music
by: Allison, Jesse, et al.
Published: (2024)
by: Allison, Jesse, et al.
Published: (2024)
MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions
by: Choi, Suhwan, et al.
Published: (2025)
by: Choi, Suhwan, et al.
Published: (2025)
Let Network Decide What to Learn: Symbolic Music Understanding Model Based on Large-scale Adversarial Pre-training
by: Zhao, Zijian
Published: (2024)
by: Zhao, Zijian
Published: (2024)
MMT-BERT: Chord-aware Symbolic Music Generation Based on Multitrack Music Transformer and MusicBERT
by: Zhu, Jinlong, et al.
Published: (2024)
by: Zhu, Jinlong, et al.
Published: (2024)
Detecting Music Performance Errors with Transformers
by: Chou, Benjamin Shiue-Hal, et al.
Published: (2025)
by: Chou, Benjamin Shiue-Hal, et al.
Published: (2025)
Tuning Music Education: AI-Powered Personalization in Learning Music
by: Sanganeria, Mayank, et al.
Published: (2024)
by: Sanganeria, Mayank, et al.
Published: (2024)
Frechet Music Distance: A Metric For Generative Symbolic Music Evaluation
by: Retkowski, Jan, et al.
Published: (2024)
by: Retkowski, Jan, et al.
Published: (2024)
PianoBART: Symbolic Piano Music Generation and Understanding with Large-Scale Pre-Training
by: Liang, Xiao, et al.
Published: (2024)
by: Liang, Xiao, et al.
Published: (2024)
Real-world Music Plagiarism Detection With Music Segment Transcription System
by: Go, Seonghyeon
Published: (2025)
by: Go, Seonghyeon
Published: (2025)
Similar Items
-
The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS
by: Carone, Brandon James, et al.
Published: (2025) -
SoundSignature: What Type of Music Do You Like?
by: Carone, Brandon James, et al.
Published: (2024) -
The Music Maestro or The Musically Challenged, A Massive Music Evaluation Benchmark for Large Language Models
by: Li, Jiajia, et al.
Published: (2024) -
Latent Acoustic Mapping for Direction of Arrival Estimation: A Self-Supervised Approach
by: Roman, Adrian S., et al.
Published: (2025) -
Content-based Controls For Music Large Language Modeling
by: Lin, Liwei, et al.
Published: (2023)