A Dataset for Named Entity Recognition and Relation Extraction from Art-historical Image Descriptions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Schneider, Stefanie, Göldl, Miriam, Stalter, Julian, Vollmer, Ricarda
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915811241754624
author Schneider, Stefanie
Göldl, Miriam
Stalter, Julian
Vollmer, Ricarda
author_facet Schneider, Stefanie
Göldl, Miriam
Stalter, Julian
Vollmer, Ricarda
contents This paper introduces FRAME (Fine-grained Recognition of Art-historical Metadata and Entities), a manually annotated dataset of art-historical image descriptions for Named Entity Recognition (NER) and Relation Extraction (RE). Descriptions were collected from museum catalogs, auction listings, open-access platforms, and scholarly databases, then filtered to ensure that each text focuses on a single artwork and contains explicit statements about its material, composition, or iconography. FRAME provides stand-off annotations in three layers: a metadata layer for object-level properties, a content layer for depicted subjects and motifs, and a co-reference layer linking repeated mentions. Across layers, entity spans are labeled with 37 types and connected by typed RE links between mentions. Entity types are aligned with Wikidata to support Named Entity Linking (NEL) and downstream knowledge-graph construction. The dataset is released as UIMA XMI Common Analysis Structure (CAS) files with accompanying images and bibliographic metadata, and can be used to benchmark and fine-tune NER and RE systems, including zero- and few-shot setups with Large Language Models (LLMs).
format Preprint
id arxiv_https___arxiv_org_abs_2602_19133
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Dataset for Named Entity Recognition and Relation Extraction from Art-historical Image Descriptions
Schneider, Stefanie
Göldl, Miriam
Stalter, Julian
Vollmer, Ricarda
Computation and Language
This paper introduces FRAME (Fine-grained Recognition of Art-historical Metadata and Entities), a manually annotated dataset of art-historical image descriptions for Named Entity Recognition (NER) and Relation Extraction (RE). Descriptions were collected from museum catalogs, auction listings, open-access platforms, and scholarly databases, then filtered to ensure that each text focuses on a single artwork and contains explicit statements about its material, composition, or iconography. FRAME provides stand-off annotations in three layers: a metadata layer for object-level properties, a content layer for depicted subjects and motifs, and a co-reference layer linking repeated mentions. Across layers, entity spans are labeled with 37 types and connected by typed RE links between mentions. Entity types are aligned with Wikidata to support Named Entity Linking (NEL) and downstream knowledge-graph construction. The dataset is released as UIMA XMI Common Analysis Structure (CAS) files with accompanying images and bibliographic metadata, and can be used to benchmark and fine-tune NER and RE systems, including zero- and few-shot setups with Large Language Models (LLMs).
title A Dataset for Named Entity Recognition and Relation Extraction from Art-historical Image Descriptions
topic Computation and Language
url https://arxiv.org/abs/2602.19133