ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Menghe, Wei, Siqing, Xing, Yuecheng, Wang, Yaheng, Meng, Fanhong, Han, Peijun, Tuan, Luu Anh, Luo, Haoran
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915949522714624
author Ma, Menghe
Wei, Siqing
Xing, Yuecheng
Wang, Yaheng
Meng, Fanhong
Han, Peijun
Tuan, Luu Anh
Luo, Haoran
author_facet Ma, Menghe
Wei, Siqing
Xing, Yuecheng
Wang, Yaheng
Meng, Fanhong
Han, Peijun
Tuan, Luu Anh
Luo, Haoran
contents Omnimodal Notation Processing (ONP) represents a unique frontier for omnimodal AI due to the rigorous, multi-dimensional alignment required across auditory, visual, and symbolic domains. Current research remains fragmented, focusing on isolated transcription tasks that fail to bridge the gap between superficial pattern recognition and the underlying musical logic. This landscape is further complicated by severe notation biases toward Western staff and the inherent unreliability of "LLM-as-a-judge" metrics, which often mask structural reasoning failures with systemic hallucinations. To establish a more rigorous standard, we introduce ONOTE, a multi-format benchmark that utilizes a deterministic pipeline--grounded in canonical pitch projection--to eliminate subjective scoring biases across diverse notation systems. Our evaluation of leading omnimodal models exposes a fundamental disconnect between perceptual accuracy and music-theoretic comprehension, providing a necessary framework for diagnosing reasoning vulnerabilities in complex, rule-constrained domains.
format Preprint
id arxiv_https___arxiv_org_abs_2604_20719
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence
Ma, Menghe
Wei, Siqing
Xing, Yuecheng
Wang, Yaheng
Meng, Fanhong
Han, Peijun
Tuan, Luu Anh
Luo, Haoran
Sound
Artificial Intelligence
Multimedia
Audio and Speech Processing
Omnimodal Notation Processing (ONP) represents a unique frontier for omnimodal AI due to the rigorous, multi-dimensional alignment required across auditory, visual, and symbolic domains. Current research remains fragmented, focusing on isolated transcription tasks that fail to bridge the gap between superficial pattern recognition and the underlying musical logic. This landscape is further complicated by severe notation biases toward Western staff and the inherent unreliability of "LLM-as-a-judge" metrics, which often mask structural reasoning failures with systemic hallucinations. To establish a more rigorous standard, we introduce ONOTE, a multi-format benchmark that utilizes a deterministic pipeline--grounded in canonical pitch projection--to eliminate subjective scoring biases across diverse notation systems. Our evaluation of leading omnimodal models exposes a fundamental disconnect between perceptual accuracy and music-theoretic comprehension, providing a necessary framework for diagnosing reasoning vulnerabilities in complex, rule-constrained domains.
title ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence
topic Sound
Artificial Intelligence
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2604.20719