Nuremberg Letterbooks: A Multi-Transcriptional Dataset of Early 15th Century Manuscripts for Document Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mayr, Martin, Krenz, Julian, Neumeier, Katharina, Bub, Anna, Bürcky, Simon, Brolich, Nina, Herbers, Klaus, Habermann, Mechthild, Fleischmann, Peter, Maier, Andreas, Christlein, Vincent
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916477190275072
author Mayr, Martin
Krenz, Julian
Neumeier, Katharina
Bub, Anna
Bürcky, Simon
Brolich, Nina
Herbers, Klaus
Habermann, Mechthild
Fleischmann, Peter
Maier, Andreas
Christlein, Vincent
author_facet Mayr, Martin
Krenz, Julian
Neumeier, Katharina
Bub, Anna
Bürcky, Simon
Brolich, Nina
Herbers, Klaus
Habermann, Mechthild
Fleischmann, Peter
Maier, Andreas
Christlein, Vincent
contents Most datasets in the field of document analysis utilize highly standardized labels, which, while simplifying specific tasks, often produce outputs that are not directly applicable to humanities research. In contrast, the Nuremberg Letterbooks dataset, which comprises historical documents from the early 15th century, addresses this gap by providing multiple types of transcriptions and accompanying metadata. This approach allows for developing methods that are more closely aligned with the needs of the humanities. The dataset includes 4 books containing 1711 labeled pages written by 10 scribes. Three types of transcriptions are provided for handwritten text recognition: Basic, diplomatic, and regularized. For the latter two, versions with and without expanded abbreviations are also available. A combination of letter ID and writer ID supports writer identification due to changing writers within pages. In the technical validation, we established baselines for various tasks, demonstrating data consistency and providing benchmarks for future research to build upon.
format Preprint
id arxiv_https___arxiv_org_abs_2411_07138
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Nuremberg Letterbooks: A Multi-Transcriptional Dataset of Early 15th Century Manuscripts for Document Analysis
Mayr, Martin
Krenz, Julian
Neumeier, Katharina
Bub, Anna
Bürcky, Simon
Brolich, Nina
Herbers, Klaus
Habermann, Mechthild
Fleischmann, Peter
Maier, Andreas
Christlein, Vincent
Computer Vision and Pattern Recognition
Most datasets in the field of document analysis utilize highly standardized labels, which, while simplifying specific tasks, often produce outputs that are not directly applicable to humanities research. In contrast, the Nuremberg Letterbooks dataset, which comprises historical documents from the early 15th century, addresses this gap by providing multiple types of transcriptions and accompanying metadata. This approach allows for developing methods that are more closely aligned with the needs of the humanities. The dataset includes 4 books containing 1711 labeled pages written by 10 scribes. Three types of transcriptions are provided for handwritten text recognition: Basic, diplomatic, and regularized. For the latter two, versions with and without expanded abbreviations are also available. A combination of letter ID and writer ID supports writer identification due to changing writers within pages. In the technical validation, we established baselines for various tasks, demonstrating data consistency and providing benchmarks for future research to build upon.
title Nuremberg Letterbooks: A Multi-Transcriptional Dataset of Early 15th Century Manuscripts for Document Analysis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.07138