Are you really listening? Boosting Perceptual Awareness in Music-QA Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Zang, Yongyi, O'Brien, Sean, Berg-Kirkpatrick, Taylor, McAuley, Julian, Novack, Zachary |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
by: Novack, Zachary, et al.
Published: (2024)
by: Novack, Zachary, et al.
Published: (2024)
PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
by: Long, Phillip, et al.
Published: (2024)
by: Long, Phillip, et al.
Published: (2024)
DITTO: Diffusion Inference-Time T-Optimization for Music Generation
by: Novack, Zachary, et al.
Published: (2024)
by: Novack, Zachary, et al.
Published: (2024)
WildFX: A DAW-Powered Pipeline for In-the-Wild Audio FX Graph Modeling
by: Yang, Qihui, et al.
Published: (2025)
by: Yang, Qihui, et al.
Published: (2025)
Steering Autoregressive Music Generation with Recursive Feature Machines
by: Zhao, Daniel, et al.
Published: (2025)
by: Zhao, Daniel, et al.
Published: (2025)
Presto! Distilling Steps and Layers for Accelerating Music Generation
by: Novack, Zachary, et al.
Published: (2024)
by: Novack, Zachary, et al.
Published: (2024)
MuseTok: Symbolic Music Tokenization for Generation and Semantic Understanding
by: Huang, Jingyue, et al.
Published: (2025)
by: Huang, Jingyue, et al.
Published: (2025)
Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset
by: Xu, Weihan, et al.
Published: (2024)
by: Xu, Weihan, et al.
Published: (2024)
WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
by: Mundada, Gagan, et al.
Published: (2025)
by: Mundada, Gagan, et al.
Published: (2025)
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
by: Kim, Haven, et al.
Published: (2025)
by: Kim, Haven, et al.
Published: (2025)
MuseCPBench: an Empirical Study of Music Editing Methods through Music Context Preservation
by: Vishe, Yash, et al.
Published: (2025)
by: Vishe, Yash, et al.
Published: (2025)
Low-Resource Guidance for Controllable Latent Audio Diffusion
by: Novack, Zachary, et al.
Published: (2026)
by: Novack, Zachary, et al.
Published: (2026)
CSyMR: Benchmarking Compositional Music Information Retrieval in Symbolic Music Reasoning
by: Wang, Boyang, et al.
Published: (2025)
by: Wang, Boyang, et al.
Published: (2025)
CoLLAP: Contrastive Long-form Language-Audio Pretraining with Musical Temporal Structure Augmentation
by: Wu, Junda, et al.
Published: (2024)
by: Wu, Junda, et al.
Published: (2024)
MusiCRS: Benchmarking Audio-Centric Conversational Recommendation
by: Surana, Rohan, et al.
Published: (2025)
by: Surana, Rohan, et al.
Published: (2025)
Reddit2Deezer: A Scalable Dataset for Real-World Grounded Conversational Music Recommendation
by: Kim, Haven, et al.
Published: (2026)
by: Kim, Haven, et al.
Published: (2026)
Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators
by: Novack, Zachary, et al.
Published: (2026)
by: Novack, Zachary, et al.
Published: (2026)
Bob's Confetti: Phonetic Memorization Attacks in Music and Video Generation
by: Roh, Jaechul, et al.
Published: (2025)
by: Roh, Jaechul, et al.
Published: (2025)
FusID: Modality-Fused Semantic IDs for Generative Music Recommendation
by: Kim, Haven, et al.
Published: (2026)
by: Kim, Haven, et al.
Published: (2026)
Composer Vector: Style-steering Symbolic Music Generation in a Latent Space
by: Jiang, Xunyi, et al.
Published: (2026)
by: Jiang, Xunyi, et al.
Published: (2026)
Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation
by: Wu, Junda, et al.
Published: (2024)
by: Wu, Junda, et al.
Published: (2024)
Voices of Civilizations: A Multilingual QA Benchmark for Global Music Understanding
by: Wu, Shangda, et al.
Published: (2026)
by: Wu, Shangda, et al.
Published: (2026)
Fast Text-to-Audio Generation with Adversarial Post-Training
by: Novack, Zachary, et al.
Published: (2025)
by: Novack, Zachary, et al.
Published: (2025)
StylePitcher: Generating Style-Following and Expressive Pitch Curves for Versatile Singing Tasks
by: Huang, Jingyue, et al.
Published: (2025)
by: Huang, Jingyue, et al.
Published: (2025)
The Interpretation Gap in Text-to-Music Generation Models
by: Zang, Yongyi, et al.
Published: (2024)
by: Zang, Yongyi, et al.
Published: (2024)
BACHI: Boundary-Aware Symbolic Chord Recognition Through Masked Iterative Decoding on Pop and Classical Music
by: Yao, Mingyang, et al.
Published: (2025)
by: Yao, Mingyang, et al.
Published: (2025)
MSRBench: A Benchmarking Dataset for Music Source Restoration
by: Zang, Yongyi, et al.
Published: (2025)
by: Zang, Yongyi, et al.
Published: (2025)
GSound-SIR: A Spatial Impulse Response Ray-Tracing and High-order Ambisonic Auralization Python Toolkit
by: Zang, Yongyi, et al.
Published: (2025)
by: Zang, Yongyi, et al.
Published: (2025)
Synthesizing Composite Hierarchical Structure from Symbolic Music Corpora
by: Shapiro, Ilana, et al.
Published: (2025)
by: Shapiro, Ilana, et al.
Published: (2025)
Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
by: Long, Phillip, et al.
Published: (2026)
by: Long, Phillip, et al.
Published: (2026)
HumMusQA: A Human-written Music Understanding QA Benchmark Dataset
by: Weck, Benno, et al.
Published: (2026)
by: Weck, Benno, et al.
Published: (2026)
Music Source Restoration
by: Zang, Yongyi, et al.
Published: (2025)
by: Zang, Yongyi, et al.
Published: (2025)
FlowSynth: Instrument Generation Through Distributional Flow Matching and Test-Time Search
by: Yang, Qihui, et al.
Published: (2025)
by: Yang, Qihui, et al.
Published: (2025)
Efficient Vocal Source Separation Through Windowed Sink Attention
by: Benetatos, Christodoulos, et al.
Published: (2025)
by: Benetatos, Christodoulos, et al.
Published: (2025)
Interpretable and Perceptually-Aligned Music Similarity with Pretrained Embeddings
by: Vohra, Arhan, et al.
Published: (2026)
by: Vohra, Arhan, et al.
Published: (2026)
Ambisonizer: Neural Upmixing as Spherical Harmonics Generation
by: Zang, Yongyi, et al.
Published: (2024)
by: Zang, Yongyi, et al.
Published: (2024)
Aligning Text-to-Music Evaluation with Human Preferences
by: Huang, Yichen, et al.
Published: (2025)
by: Huang, Yichen, et al.
Published: (2025)
SAME: A Semantically-Aligned Music Autoencoder
by: Parker, Julian D., et al.
Published: (2026)
by: Parker, Julian D., et al.
Published: (2026)
Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering
by: Koh, Junyoung, et al.
Published: (2026)
by: Koh, Junyoung, et al.
Published: (2026)
Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
by: Paek, Nathan, et al.
Published: (2025)
by: Paek, Nathan, et al.
Published: (2025)
Similar Items
-
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
by: Novack, Zachary, et al.
Published: (2024) -
PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
by: Long, Phillip, et al.
Published: (2024) -
DITTO: Diffusion Inference-Time T-Optimization for Music Generation
by: Novack, Zachary, et al.
Published: (2024) -
WildFX: A DAW-Powered Pipeline for In-the-Wild Audio FX Graph Modeling
by: Yang, Qihui, et al.
Published: (2025) -
Steering Autoregressive Music Generation with Recursive Feature Machines
by: Zhao, Daniel, et al.
Published: (2025)