CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Yinghao, Xia, Haiwen, Gao, Hewei, Chen, Weixiong, Ye, Yuxin, Yang, Yuchen, Chang, Sungkyun, Ding, Mingshuo, Li, Yizhi, Yuan, Ruibin, Dixon, Simon, Benetos, Emmanouil |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
von: Ma, Yinghao, et al.
Veröffentlicht: (2025)
von: Ma, Yinghao, et al.
Veröffentlicht: (2025)
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
von: Zhou, Ziya, et al.
Veröffentlicht: (2024)
von: Zhou, Ziya, et al.
Veröffentlicht: (2024)
RUMAA: Repeat-Aware Unified Music Audio Analysis for Score-Performance Alignment, Transcription, and Mistake Detection
von: Chang, Sungkyun, et al.
Veröffentlicht: (2025)
von: Chang, Sungkyun, et al.
Veröffentlicht: (2025)
AutoMV: An Automatic Multi-Agent System for Music Video Generation
von: Tang, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Tang, Xiaoxuan, et al.
Veröffentlicht: (2025)
YourMT3+: Multi-instrument Music Transcription with Enhanced Transformer Architectures and Cross-dataset Stem Augmentation
von: Chang, Sungkyun, et al.
Veröffentlicht: (2024)
von: Chang, Sungkyun, et al.
Veröffentlicht: (2024)
MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
von: Weck, Benno, et al.
Veröffentlicht: (2024)
von: Weck, Benno, et al.
Veröffentlicht: (2024)
MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response
von: Deng, Zihao, et al.
Veröffentlicht: (2023)
von: Deng, Zihao, et al.
Veröffentlicht: (2023)
ComposerX: Multi-Agent Symbolic Music Composition with LLMs
von: Deng, Qixin, et al.
Veröffentlicht: (2024)
von: Deng, Qixin, et al.
Veröffentlicht: (2024)
A Survey on Multimodal Music Emotion Recognition
von: Liyanarachchi, Rashini, et al.
Veröffentlicht: (2025)
von: Liyanarachchi, Rashini, et al.
Veröffentlicht: (2025)
Dance-to-Music Generation with Encoder-based Textual Inversion
von: Li, Sifei, et al.
Veröffentlicht: (2024)
von: Li, Sifei, et al.
Veröffentlicht: (2024)
A Survey on Cross-Modal Interaction Between Music and Multimodal Data
von: Li, Sifei, et al.
Veröffentlicht: (2025)
von: Li, Sifei, et al.
Veröffentlicht: (2025)
MusFlow: Multimodal Music Generation via Conditional Flow Matching
von: Song, Jiahao, et al.
Veröffentlicht: (2025)
von: Song, Jiahao, et al.
Veröffentlicht: (2025)
MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
von: Zhao, Qihao, et al.
Veröffentlicht: (2026)
von: Zhao, Qihao, et al.
Veröffentlicht: (2026)
MeloTrans: A Text to Symbolic Music Generation Model Following Human Composition Habit
von: Wang, Yutian, et al.
Veröffentlicht: (2024)
von: Wang, Yutian, et al.
Veröffentlicht: (2024)
MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
Flexible Control in Symbolic Music Generation via Musical Metadata
von: Han, Sangjun, et al.
Veröffentlicht: (2024)
von: Han, Sangjun, et al.
Veröffentlicht: (2024)
MusicSem: A Semantically Rich Language--Audio Dataset of Natural Music Descriptions
von: Salganik, Rebecca, et al.
Veröffentlicht: (2026)
von: Salganik, Rebecca, et al.
Veröffentlicht: (2026)
Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
MusicAOG: an Energy-Based Model for Learning and Sampling a Hierarchical Representation of Symbolic Music
von: Qian, Yikai, et al.
Veröffentlicht: (2024)
von: Qian, Yikai, et al.
Veröffentlicht: (2024)
Intelligent Text-Conditioned Music Generation
von: Xie, Zhouyao, et al.
Veröffentlicht: (2024)
von: Xie, Zhouyao, et al.
Veröffentlicht: (2024)
Optimizing Feature Extraction for Symbolic Music
von: Simonetta, Federico, et al.
Veröffentlicht: (2023)
von: Simonetta, Federico, et al.
Veröffentlicht: (2023)
YuE: Scaling Open Foundation Models for Long-Form Music Generation
von: Yuan, Ruibin, et al.
Veröffentlicht: (2025)
von: Yuan, Ruibin, et al.
Veröffentlicht: (2025)
MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
A Survey on Evaluation Metrics for Music Generation
von: Kader, Faria Binte, et al.
Veröffentlicht: (2025)
von: Kader, Faria Binte, et al.
Veröffentlicht: (2025)
Emotion-Aligned Contrastive Learning Between Images and Music
von: Stewart, Shanti, et al.
Veröffentlicht: (2023)
von: Stewart, Shanti, et al.
Veröffentlicht: (2023)
Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tuning
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval
von: Stewart, Shanti, et al.
Veröffentlicht: (2024)
von: Stewart, Shanti, et al.
Veröffentlicht: (2024)
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
ScripTONES: Sentiment-Conditioned Music Generation for Movie Scripts
von: Veerendranath, Vishruth, et al.
Veröffentlicht: (2024)
von: Veerendranath, Vishruth, et al.
Veröffentlicht: (2024)
MART: Learning Hierarchical Music Audio Representations with Part-Whole Transformer
von: Yao, Dong, et al.
Veröffentlicht: (2023)
von: Yao, Dong, et al.
Veröffentlicht: (2023)
Jamendo-QA: A Large-Scale Music Question Answering Dataset
von: Koh, Junyoung, et al.
Veröffentlicht: (2025)
von: Koh, Junyoung, et al.
Veröffentlicht: (2025)
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
von: Kim, Haven, et al.
Veröffentlicht: (2025)
von: Kim, Haven, et al.
Veröffentlicht: (2025)
MoMu-Diffusion: On Learning Long-Term Motion-Music Synchronization and Correspondence
von: You, Fuming, et al.
Veröffentlicht: (2024)
von: You, Fuming, et al.
Veröffentlicht: (2024)
SONIQUE: Video Background Music Generation Using Unpaired Audio-Visual Data
von: Zhang, Liqian, et al.
Veröffentlicht: (2024)
von: Zhang, Liqian, et al.
Veröffentlicht: (2024)
Enhancing Expressiveness in Dance Generation via Integrating Frequency and Music Style Information
von: Huang, Qiaochu, et al.
Veröffentlicht: (2024)
von: Huang, Qiaochu, et al.
Veröffentlicht: (2024)
MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions
von: Choi, Suhwan, et al.
Veröffentlicht: (2025)
von: Choi, Suhwan, et al.
Veröffentlicht: (2025)
WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
von: Ji, Shengpeng, et al.
Veröffentlicht: (2025)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2025)
Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion Modeling
von: Li, Xiaojie, et al.
Veröffentlicht: (2025)
von: Li, Xiaojie, et al.
Veröffentlicht: (2025)
CoheDancers: Enhancing Interactive Group Dance Generation through Music-Driven Coherence Decomposition
von: Yang, Kaixing, et al.
Veröffentlicht: (2024)
von: Yang, Kaixing, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
von: Ma, Yinghao, et al.
Veröffentlicht: (2025) -
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
von: Zhou, Ziya, et al.
Veröffentlicht: (2024) -
RUMAA: Repeat-Aware Unified Music Audio Analysis for Score-Performance Alignment, Transcription, and Mistake Detection
von: Chang, Sungkyun, et al.
Veröffentlicht: (2025) -
AutoMV: An Automatic Multi-Agent System for Music Video Generation
von: Tang, Xiaoxuan, et al.
Veröffentlicht: (2025) -
YourMT3+: Multi-instrument Music Transcription with Enhanced Transformer Architectures and Cross-dataset Stem Augmentation
von: Chang, Sungkyun, et al.
Veröffentlicht: (2024)