GaMMA: Towards Joint Global-Temporal Music Understanding in Large Multimodal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | You, Zuyao, Yu, Zhesong, Liu, Mingyu, Zhu, Bilei, Wan, Yuan, Wu, Zuxuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MINT: Boosting Audio-Language Model via Multi-Target Pre-Training and Instruction Tuning
von: Zhao, Hang, et al.
Veröffentlicht: (2024)
von: Zhao, Hang, et al.
Veröffentlicht: (2024)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
von: Shi, Jiapeng, et al.
Veröffentlicht: (2026)
von: Shi, Jiapeng, et al.
Veröffentlicht: (2026)
ByteComposer: a Human-like Melody Composition Method based on Language Model Agent
von: Liang, Xia, et al.
Veröffentlicht: (2024)
von: Liang, Xia, et al.
Veröffentlicht: (2024)
Voices of Civilizations: A Multilingual QA Benchmark for Global Music Understanding
von: Wu, Shangda, et al.
Veröffentlicht: (2026)
von: Wu, Shangda, et al.
Veröffentlicht: (2026)
Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
von: Tong, Xinyi, et al.
Veröffentlicht: (2025)
von: Tong, Xinyi, et al.
Veröffentlicht: (2025)
Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical Scores
von: Dai, Congren, et al.
Veröffentlicht: (2025)
von: Dai, Congren, et al.
Veröffentlicht: (2025)
PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music
von: Bang, Hayeon, et al.
Veröffentlicht: (2025)
von: Bang, Hayeon, et al.
Veröffentlicht: (2025)
Towards Effective Negation Modeling in Joint Audio-Text Models for Music
von: Vasilakis, Yannis, et al.
Veröffentlicht: (2026)
von: Vasilakis, Yannis, et al.
Veröffentlicht: (2026)
Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning
von: You, Zuyao, et al.
Veröffentlicht: (2025)
von: You, Zuyao, et al.
Veröffentlicht: (2025)
Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
von: Tal, Or, et al.
Veröffentlicht: (2024)
von: Tal, Or, et al.
Veröffentlicht: (2024)
The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models
von: You, Yuhuan, et al.
Veröffentlicht: (2026)
von: You, Yuhuan, et al.
Veröffentlicht: (2026)
MusicLIME: Explainable Multimodal Music Understanding
von: Sotirou, Theodoros, et al.
Veröffentlicht: (2024)
von: Sotirou, Theodoros, et al.
Veröffentlicht: (2024)
Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation
von: Wu, Junda, et al.
Veröffentlicht: (2024)
von: Wu, Junda, et al.
Veröffentlicht: (2024)
MIDI-LLaMA: An Instruction-Following Multimodal LLM for Symbolic Music Understanding
von: Yang, Meng, et al.
Veröffentlicht: (2026)
von: Yang, Meng, et al.
Veröffentlicht: (2026)
M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
von: Liu, Shansong, et al.
Veröffentlicht: (2023)
von: Liu, Shansong, et al.
Veröffentlicht: (2023)
MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models
von: Liu, Shansong, et al.
Veröffentlicht: (2024)
von: Liu, Shansong, et al.
Veröffentlicht: (2024)
Towards Practical Real-Time Low-Latency Music Source Separation
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
TALKPLAY: Multimodal Music Recommendation with Large Language Models
von: Doh, Seungheon, et al.
Veröffentlicht: (2025)
von: Doh, Seungheon, et al.
Veröffentlicht: (2025)
ABC-Eval: Benchmarking Large Language Models on Symbolic Music Understanding and Instruction Following
von: Zhao, Jiahao, et al.
Veröffentlicht: (2025)
von: Zhao, Jiahao, et al.
Veröffentlicht: (2025)
TinyMU: A Compact Audio-Language Model for Music Understanding
von: Li, Xiquan, et al.
Veröffentlicht: (2026)
von: Li, Xiquan, et al.
Veröffentlicht: (2026)
CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models
von: Wu, Shangda, et al.
Veröffentlicht: (2024)
von: Wu, Shangda, et al.
Veröffentlicht: (2024)
WeaveMuse: An Open Agentic System for Multimodal Music Understanding and Generation
von: Karystinaios, Emmanouil
Veröffentlicht: (2025)
von: Karystinaios, Emmanouil
Veröffentlicht: (2025)
MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
von: Weck, Benno, et al.
Veröffentlicht: (2024)
von: Weck, Benno, et al.
Veröffentlicht: (2024)
NOTA: Multimodal Music Notation Understanding for Visual Large Language Model
von: Tang, Mingni, et al.
Veröffentlicht: (2025)
von: Tang, Mingni, et al.
Veröffentlicht: (2025)
Versatile Symbolic Music-for-Music Modeling via Function Alignment
von: Jiang, Junyan, et al.
Veröffentlicht: (2025)
von: Jiang, Junyan, et al.
Veröffentlicht: (2025)
SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing
von: Ma, Ziyang, et al.
Veröffentlicht: (2026)
von: Ma, Ziyang, et al.
Veröffentlicht: (2026)
Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
von: Sun, Haoqin, et al.
Veröffentlicht: (2026)
von: Sun, Haoqin, et al.
Veröffentlicht: (2026)
Khala: Scaling Acoustic Token Language Models Toward High-Fidelity Music Generation
von: Liu, Jiafeng, et al.
Veröffentlicht: (2026)
von: Liu, Jiafeng, et al.
Veröffentlicht: (2026)
Evaluating Multimodal Large Language Models on Core Music Perception Tasks
von: Carone, Brandon James, et al.
Veröffentlicht: (2025)
von: Carone, Brandon James, et al.
Veröffentlicht: (2025)
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
von: Zhou, Ziya, et al.
Veröffentlicht: (2024)
von: Zhou, Ziya, et al.
Veröffentlicht: (2024)
MQAD: A Large-Scale Question Answering Dataset for Training Music Large Language Models
von: Ouyang, Zhihao, et al.
Veröffentlicht: (2025)
von: Ouyang, Zhihao, et al.
Veröffentlicht: (2025)
DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2025)
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2025)
HeartMuLa: A Family of Open Sourced Music Foundation Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2026)
von: Yang, Dongchao, et al.
Veröffentlicht: (2026)
A Survey of Foundation Models for Music Understanding
von: Li, Wenjun, et al.
Veröffentlicht: (2024)
von: Li, Wenjun, et al.
Veröffentlicht: (2024)
Synthesize, Diagnose, and Optimize: Towards Fine-Grained Vision-Language Understanding
von: Peng, Wujian, et al.
Veröffentlicht: (2023)
von: Peng, Wujian, et al.
Veröffentlicht: (2023)
Improving BERT for Symbolic Music Understanding Using Token Denoising and Pianoroll Prediction
von: Wang, Jun-You, et al.
Veröffentlicht: (2025)
von: Wang, Jun-You, et al.
Veröffentlicht: (2025)
BeatFM: Improving Beat Tracking with Pre-trained Music Foundation Model
von: Ru, Ganghui, et al.
Veröffentlicht: (2025)
von: Ru, Ganghui, et al.
Veröffentlicht: (2025)
MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
von: Zhao, Qihao, et al.
Veröffentlicht: (2026)
von: Zhao, Qihao, et al.
Veröffentlicht: (2026)
Who Will Top the Charts? Multimodal Music Popularity Prediction via Adaptive Fusion of Modality Experts and Temporal Engagement Modeling
von: Choudhary, Yash, et al.
Veröffentlicht: (2025)
von: Choudhary, Yash, et al.
Veröffentlicht: (2025)
MAJL: A Model-Agnostic Joint Learning Framework for Music Source Separation and Pitch Estimation
von: Wei, Haojie, et al.
Veröffentlicht: (2025)
von: Wei, Haojie, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MINT: Boosting Audio-Language Model via Multi-Target Pre-Training and Instruction Tuning
von: Zhao, Hang, et al.
Veröffentlicht: (2024) -
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
von: Shi, Jiapeng, et al.
Veröffentlicht: (2026) -
ByteComposer: a Human-like Melody Composition Method based on Language Model Agent
von: Liang, Xia, et al.
Veröffentlicht: (2024) -
Voices of Civilizations: A Multilingual QA Benchmark for Global Music Understanding
von: Wu, Shangda, et al.
Veröffentlicht: (2026) -
Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
von: Tong, Xinyi, et al.
Veröffentlicht: (2025)