AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kišš, Martin, Hradiš, Michal, Dvořáková, Martina, Jiroušek, Václav, Kersch, Filip
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916845861208064
author Kišš, Martin
Hradiš, Michal
Dvořáková, Martina
Jiroušek, Václav
Kersch, Filip
author_facet Kišš, Martin
Hradiš, Michal
Dvořáková, Martina
Jiroušek, Václav
Kersch, Filip
contents We introduce the AnnoPage Dataset, a novel collection of 7,550 pages from historical documents, primarily in Czech and German, spanning from 1485 to the present, focusing on the late 19th and early 20th centuries. The dataset is designed to support research in document layout analysis and object detection. Each page is annotated with axis-aligned bounding boxes (AABB) representing elements of 25 categories of non-textual elements, such as images, maps, decorative elements, or charts, following the Czech Methodology of image document processing. The annotations were created by expert librarians to ensure accuracy and consistency. The dataset also incorporates pages from multiple, mainly historical, document datasets to enhance variability and maintain continuity. The dataset is divided into development and test subsets, with the test set carefully selected to maintain the category distribution. We provide baseline results using YOLO and DETR object detectors, offering a reference point for future research. The AnnoPage Dataset is publicly available on Zenodo (https://doi.org/10.5281/zenodo.12788419), along with ground-truth annotations in YOLO format.
format Preprint
id arxiv_https___arxiv_org_abs_2503_22526
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization
Kišš, Martin
Hradiš, Michal
Dvořáková, Martina
Jiroušek, Václav
Kersch, Filip
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
We introduce the AnnoPage Dataset, a novel collection of 7,550 pages from historical documents, primarily in Czech and German, spanning from 1485 to the present, focusing on the late 19th and early 20th centuries. The dataset is designed to support research in document layout analysis and object detection. Each page is annotated with axis-aligned bounding boxes (AABB) representing elements of 25 categories of non-textual elements, such as images, maps, decorative elements, or charts, following the Czech Methodology of image document processing. The annotations were created by expert librarians to ensure accuracy and consistency. The dataset also incorporates pages from multiple, mainly historical, document datasets to enhance variability and maintain continuity. The dataset is divided into development and test subsets, with the test set carefully selected to maintain the category distribution. We provide baseline results using YOLO and DETR object detectors, offering a reference point for future research. The AnnoPage Dataset is publicly available on Zenodo (https://doi.org/10.5281/zenodo.12788419), along with ground-truth annotations in YOLO format.
title AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.22526