Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering
Fuente:
arXiv
Guardado en:
| Autores principales: | Koh, Junyoung, Lee, Jaeyun, Kim, Soo Yong, Choi, Gyu Hyeong, Koh, Jung In, Phillips, Jordan, Lee, Yeonjin, Song, Min |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Jamendo-QA: A Large-Scale Music Question Answering Dataset
por: Koh, Junyoung, et al.
Publicado: (2025)
por: Koh, Junyoung, et al.
Publicado: (2025)
PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music
por: Bang, Hayeon, et al.
Publicado: (2025)
por: Bang, Hayeon, et al.
Publicado: (2025)
Music4All A+A: A Multimodal Dataset for Music Information Retrieval Tasks
por: Geiger, Jonas, et al.
Publicado: (2025)
por: Geiger, Jonas, et al.
Publicado: (2025)
Enriching Music Descriptions with a Finetuned-LLM and Metadata for Text-to-Music Retrieval
por: Doh, SeungHeon, et al.
Publicado: (2024)
por: Doh, SeungHeon, et al.
Publicado: (2024)
AIBA: Attention-based Instrument Band Alignment for Text-to-Audio Diffusion
por: Koh, Junyoung, et al.
Publicado: (2025)
por: Koh, Junyoung, et al.
Publicado: (2025)
TalkPlay-Tools: Conversational Music Recommendation with LLM Tool Calling
por: Doh, Seungheon, et al.
Publicado: (2025)
por: Doh, Seungheon, et al.
Publicado: (2025)
Diff4Steer: Steerable Diffusion Prior for Generative Music Retrieval with Semantic Guidance
por: Bao, Xuchan, et al.
Publicado: (2024)
por: Bao, Xuchan, et al.
Publicado: (2024)
EMO100DB: An Open Dataset of Improvised Songs with Emotion Data
por: Hwang, Daeun, et al.
Publicado: (2025)
por: Hwang, Daeun, et al.
Publicado: (2025)
OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset
por: Tran, Quang-Linh, et al.
Publicado: (2025)
por: Tran, Quang-Linh, et al.
Publicado: (2025)
TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation
por: Choi, Keunwoo, et al.
Publicado: (2025)
por: Choi, Keunwoo, et al.
Publicado: (2025)
JEPOO: Highly Accurate Joint Estimation of Pitch, Onset and Offset for Music Information Retrieval
por: Wei, Haojie, et al.
Publicado: (2023)
por: Wei, Haojie, et al.
Publicado: (2023)
Instrumental Text-to-Music Generation with Auxiliary Conditioning Branches
por: Koh, Junyoung
Publicado: (2026)
por: Koh, Junyoung
Publicado: (2026)
Learning Normal Patterns in Musical Loops
por: Dadman, Shayan, et al.
Publicado: (2025)
por: Dadman, Shayan, et al.
Publicado: (2025)
A Dataset and Baselines for Measuring and Predicting the Music Piece Memorability
por: Tseng, Li-Yang, et al.
Publicado: (2024)
por: Tseng, Li-Yang, et al.
Publicado: (2024)
Music Genre Classification: Ensemble Learning with Subcomponents-level Attention
por: Liu, Yichen, et al.
Publicado: (2024)
por: Liu, Yichen, et al.
Publicado: (2024)
MERGE -- A Bimodal Audio-Lyrics Dataset for Static Music Emotion Recognition
por: Louro, Pedro Lima, et al.
Publicado: (2024)
por: Louro, Pedro Lima, et al.
Publicado: (2024)
Enhancing Automatic Chord Recognition via Pseudo-Labeling and Knowledge Distillation
por: Phan, Nghia, et al.
Publicado: (2026)
por: Phan, Nghia, et al.
Publicado: (2026)
On the Effect of Data-Augmentation on Local Embedding Properties in the Contrastive Learning of Music Audio Representations
por: McCallum, Matthew C., et al.
Publicado: (2024)
por: McCallum, Matthew C., et al.
Publicado: (2024)
Learning Musical Representations for Music Performance Question Answering
por: Diao, Xingjian, et al.
Publicado: (2025)
por: Diao, Xingjian, et al.
Publicado: (2025)
Streaming Piano Transcription Based on Consistent Onset and Offset Decoding with Sustain Pedal Detection
por: Wei, Weixing, et al.
Publicado: (2025)
por: Wei, Weixing, et al.
Publicado: (2025)
Flexible Control in Symbolic Music Generation via Musical Metadata
por: Han, Sangjun, et al.
Publicado: (2024)
por: Han, Sangjun, et al.
Publicado: (2024)
Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
por: Tong, Xinyi, et al.
Publicado: (2025)
por: Tong, Xinyi, et al.
Publicado: (2025)
MusicWeaver: Composer-Style Structural Editing and Minute-Scale Coherent Music Generation
por: Wang, Xuanchen, et al.
Publicado: (2025)
por: Wang, Xuanchen, et al.
Publicado: (2025)
Learning Sparsity for Effective and Efficient Music Performance Question Answering
por: Diao, Xingjian, et al.
Publicado: (2025)
por: Diao, Xingjian, et al.
Publicado: (2025)
Gesture2Music: A Low-Latency Real-Time Framework for Continuous Gesture-Driven Music Generation
por: Jeyaraj, Rathinaraja, et al.
Publicado: (2025)
por: Jeyaraj, Rathinaraja, et al.
Publicado: (2025)
Learning Audio-Visual Embeddings with Inferred Latent Interaction Graphs
por: Zeng, Donghuo, et al.
Publicado: (2026)
por: Zeng, Donghuo, et al.
Publicado: (2026)
Music Audio-Visual Question Answering Requires Specialized Multimodal Designs
por: You, Wenhao, et al.
Publicado: (2025)
por: You, Wenhao, et al.
Publicado: (2025)
Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model
por: Karchkhadze, Tornike, et al.
Publicado: (2024)
por: Karchkhadze, Tornike, et al.
Publicado: (2024)
MR-MT3: Memory Retaining Multi-Track Music Transcription to Mitigate Instrument Leakage
por: Tan, Hao Hao, et al.
Publicado: (2024)
por: Tan, Hao Hao, et al.
Publicado: (2024)
HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation
por: Zhu, Jian, et al.
Publicado: (2026)
por: Zhu, Jian, et al.
Publicado: (2026)
TAGF: Time-aware Gated Fusion for Multimodal Valence-Arousal Estimation
por: Lee, Yubeen, et al.
Publicado: (2025)
por: Lee, Yubeen, et al.
Publicado: (2025)
MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning
por: Gong, Ziyu, et al.
Publicado: (2025)
por: Gong, Ziyu, et al.
Publicado: (2025)
Towards Practical Real-Time Low-Latency Music Source Separation
por: Wu, Junyu, et al.
Publicado: (2025)
por: Wu, Junyu, et al.
Publicado: (2025)
MIDI-LLaMA: An Instruction-Following Multimodal LLM for Symbolic Music Understanding
por: Yang, Meng, et al.
Publicado: (2026)
por: Yang, Meng, et al.
Publicado: (2026)
Anchor-aware Deep Metric Learning for Audio-visual Retrieval
por: Zeng, Donghuo, et al.
Publicado: (2024)
por: Zeng, Donghuo, et al.
Publicado: (2024)
ASK: Adaptive Self-improving Knowledge Framework for Audio Text Retrieval
por: Fu, Siyuan, et al.
Publicado: (2025)
por: Fu, Siyuan, et al.
Publicado: (2025)
SynthTab: Leveraging Synthesized Data for Guitar Tablature Transcription
por: Zang, Yongyi, et al.
Publicado: (2023)
por: Zang, Yongyi, et al.
Publicado: (2023)
MG-Former: A Transformer-Based Framework for Music-Driven 3D Conducting Gesture Generation
por: Qiu, Ke, et al.
Publicado: (2026)
por: Qiu, Ke, et al.
Publicado: (2026)
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
por: Niu, Xinlei, et al.
Publicado: (2025)
por: Niu, Xinlei, et al.
Publicado: (2025)
Music Arena: Live Evaluation for Text-to-Music
por: Kim, Yonghyun, et al.
Publicado: (2025)
por: Kim, Yonghyun, et al.
Publicado: (2025)
Ejemplares similares
-
Jamendo-QA: A Large-Scale Music Question Answering Dataset
por: Koh, Junyoung, et al.
Publicado: (2025) -
PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music
por: Bang, Hayeon, et al.
Publicado: (2025) -
Music4All A+A: A Multimodal Dataset for Music Information Retrieval Tasks
por: Geiger, Jonas, et al.
Publicado: (2025) -
Enriching Music Descriptions with a Finetuned-LLM and Metadata for Text-to-Music Retrieval
por: Doh, SeungHeon, et al.
Publicado: (2024) -
AIBA: Attention-based Instrument Band Alignment for Text-to-Audio Diffusion
por: Koh, Junyoung, et al.
Publicado: (2025)