Voices of Civilizations: A Multilingual QA Benchmark for Global Music Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Shangda, Zhou, Ziya, Zang, Yongyi, Zheng, Yutong, Liang, Dafang, Yuan, Ruibin, Kong, Qiuqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Music Source Restoration
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
Training-Free Multi-Step Audio Source Separation
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
Piano Transcription by Hierarchical Language Modeling with Pretrained Roll-based Encoders
von: Li, Dichucheng, et al.
Veröffentlicht: (2025)
von: Li, Dichucheng, et al.
Veröffentlicht: (2025)
Summary of The Inaugural Music Source Restoration Challenge
von: Zang, Yongyi, et al.
Veröffentlicht: (2026)
von: Zang, Yongyi, et al.
Veröffentlicht: (2026)
MusicScore: A Dataset for Music Score Modeling and Generation
von: Lin, Yuheng, et al.
Veröffentlicht: (2024)
von: Lin, Yuheng, et al.
Veröffentlicht: (2024)
The Interpretation Gap in Text-to-Music Generation Models
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
von: Zhou, Ziya, et al.
Veröffentlicht: (2024)
von: Zhou, Ziya, et al.
Veröffentlicht: (2024)
CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages
von: Wu, Shangda, et al.
Veröffentlicht: (2025)
von: Wu, Shangda, et al.
Veröffentlicht: (2025)
MelodyT5: A Unified Score-to-Score Transformer for Symbolic Music Processing
von: Wu, Shangda, et al.
Veröffentlicht: (2024)
von: Wu, Shangda, et al.
Veröffentlicht: (2024)
Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2023)
von: Wang, Zhichao, et al.
Veröffentlicht: (2023)
SingFake: Singing Voice Deepfake Detection
von: Zang, Yongyi, et al.
Veröffentlicht: (2023)
von: Zang, Yongyi, et al.
Veröffentlicht: (2023)
Exploring Tokenization Methods for Multitrack Sheet Music Generation
von: Wang, Yashan, et al.
Veröffentlicht: (2024)
von: Wang, Yashan, et al.
Veröffentlicht: (2024)
Ambisonizer: Neural Upmixing as Spherical Harmonics Generation
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
SVDD 2024: The Inaugural Singing Voice Deepfake Detection Challenge
von: Zhang, You, et al.
Veröffentlicht: (2024)
von: Zhang, You, et al.
Veröffentlicht: (2024)
SymPAC: Scalable Symbolic Music Generation With Prompts And Constraints
von: Chen, Haonan, et al.
Veröffentlicht: (2024)
von: Chen, Haonan, et al.
Veröffentlicht: (2024)
CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
MuPT: A Generative Symbolic Music Pretrained Transformer
von: Qu, Xingwei, et al.
Veröffentlicht: (2024)
von: Qu, Xingwei, et al.
Veröffentlicht: (2024)
Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer
von: Hou, Siyuan, et al.
Veröffentlicht: (2024)
von: Hou, Siyuan, et al.
Veröffentlicht: (2024)
CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models
von: Wu, Shangda, et al.
Veröffentlicht: (2024)
von: Wu, Shangda, et al.
Veröffentlicht: (2024)
TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
von: Farhadipour, Aref, et al.
Veröffentlicht: (2026)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2026)
Selective-Memory Meta-Learning with Environment Representations for Sound Event Localization and Detection
von: Hu, Jinbo, et al.
Veröffentlicht: (2023)
von: Hu, Jinbo, et al.
Veröffentlicht: (2023)
NotaGen: Advancing Musicality in Symbolic Music Generation with Large Language Model Training Paradigms
von: Wang, Yashan, et al.
Veröffentlicht: (2025)
von: Wang, Yashan, et al.
Veröffentlicht: (2025)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion
von: Ferreira, Alef Iury Siqueira, et al.
Veröffentlicht: (2025)
von: Ferreira, Alef Iury Siqueira, et al.
Veröffentlicht: (2025)
SemanticAudio: Audio Generation and Editing in Semantic Space
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
Language-Queried Target Sound Extraction Without Parallel Training Data
von: Ma, Hao, et al.
Veröffentlicht: (2024)
von: Ma, Hao, et al.
Veröffentlicht: (2024)
Jamendo-QA: A Large-Scale Music Question Answering Dataset
von: Koh, Junyoung, et al.
Veröffentlicht: (2025)
von: Koh, Junyoung, et al.
Veröffentlicht: (2025)
PianoBART: Symbolic Piano Music Generation and Understanding with Large-Scale Pre-Training
von: Liang, Xiao, et al.
Veröffentlicht: (2024)
von: Liang, Xiao, et al.
Veröffentlicht: (2024)
MelTok: 2D Tokenization for Single-Codebook Audio Compression
von: Li, Jingyi, et al.
Veröffentlicht: (2025)
von: Li, Jingyi, et al.
Veröffentlicht: (2025)
YingMusic-Singer-Plus: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance
von: Hao, Chunbo, et al.
Veröffentlicht: (2026)
von: Hao, Chunbo, et al.
Veröffentlicht: (2026)
A Holistic Evaluation of Piano Sound Quality
von: Zhou, Monan, et al.
Veröffentlicht: (2023)
von: Zhou, Monan, et al.
Veröffentlicht: (2023)
SingVERSE: A Diverse, Real-World Benchmark for Singing Voice Enhancement
von: Jiang, Shaohan, et al.
Veröffentlicht: (2025)
von: Jiang, Shaohan, et al.
Veröffentlicht: (2025)
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
von: Hou, Yixuan, et al.
Veröffentlicht: (2025)
von: Hou, Yixuan, et al.
Veröffentlicht: (2025)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
von: Chen, Huakang, et al.
Veröffentlicht: (2026)
von: Chen, Huakang, et al.
Veröffentlicht: (2026)
DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning
von: Li, Xiquan, et al.
Veröffentlicht: (2024)
von: Li, Xiquan, et al.
Veröffentlicht: (2024)
EchoVoices: Preserving Generational Voices and Memories for Seniors and Children
von: Xu, Haiying, et al.
Veröffentlicht: (2025)
von: Xu, Haiying, et al.
Veröffentlicht: (2025)
SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge Evaluation Plan
von: Zhang, You, et al.
Veröffentlicht: (2024)
von: Zhang, You, et al.
Veröffentlicht: (2024)
SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Music Source Restoration
von: Zang, Yongyi, et al.
Veröffentlicht: (2025) -
Training-Free Multi-Step Audio Source Separation
von: Zang, Yongyi, et al.
Veröffentlicht: (2025) -
Piano Transcription by Hierarchical Language Modeling with Pretrained Roll-based Encoders
von: Li, Dichucheng, et al.
Veröffentlicht: (2025) -
Summary of The Inaugural Music Source Restoration Challenge
von: Zang, Yongyi, et al.
Veröffentlicht: (2026) -
MusicScore: A Dataset for Music Score Modeling and Generation
von: Lin, Yuheng, et al.
Veröffentlicht: (2024)