Direct content-based retrieval from music scores images
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913170435604480 |
|---|---|
| author | Luna-Barahona, Noelia Ríos-Vila, Antonio Fuentes-Hurtado, Félix Rizo, David Calvo-Zaragoza, Jorge |
| author_facet | Luna-Barahona, Noelia Ríos-Vila, Antonio Fuentes-Hurtado, Félix Rizo, David Calvo-Zaragoza, Jorge |
| contents | The digitization of musical scores plays a crucial role in their preservation and accessibility, yet information retrieval still depends mainly on metadata searches, such as by title or composer. Content based search in music score images remains underexplored compared to text documents, despite its potential value for musicians, musicologists, and educators. This work contributes to the field by first studying which characteristics of a score are most relevant for search and by defining a systematic method to build query datasets from any annotated corpus. We also consider diverse methods for content-based search on music score images, ranging from transcription-based approaches relying on Optical Music Recognition (OMR), to a transcription-free Transformer model trained to recognize queries directly from score images, and a text-prompted Large Language Model. Our experiments evaluate these models on four corpora exhibiting diverse characteristics in terms of dataset size, image quality, and typesetting mechanisms. Overall, each method excels under different conditions: OMR-based pipelines achieve higher in-domain retrieval, whereas transcription-free models handle domain variability more effectively. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_22255 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Direct content-based retrieval from music scores images Luna-Barahona, Noelia Ríos-Vila, Antonio Fuentes-Hurtado, Félix Rizo, David Calvo-Zaragoza, Jorge Computer Vision and Pattern Recognition Information Retrieval The digitization of musical scores plays a crucial role in their preservation and accessibility, yet information retrieval still depends mainly on metadata searches, such as by title or composer. Content based search in music score images remains underexplored compared to text documents, despite its potential value for musicians, musicologists, and educators. This work contributes to the field by first studying which characteristics of a score are most relevant for search and by defining a systematic method to build query datasets from any annotated corpus. We also consider diverse methods for content-based search on music score images, ranging from transcription-based approaches relying on Optical Music Recognition (OMR), to a transcription-free Transformer model trained to recognize queries directly from score images, and a text-prompted Large Language Model. Our experiments evaluate these models on four corpora exhibiting diverse characteristics in terms of dataset size, image quality, and typesetting mechanisms. Overall, each method excels under different conditions: OMR-based pipelines achieve higher in-domain retrieval, whereas transcription-free models handle domain variability more effectively. |
| title | Direct content-based retrieval from music scores images |
| topic | Computer Vision and Pattern Recognition Information Retrieval |
| url | https://arxiv.org/abs/2605.22255 |