TRIDIS: A Comprehensive Medieval and Early Modern Corpus for HTR and NER

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Aguilar, Sergio Torres
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912332978847744
author Aguilar, Sergio Torres
author_facet Aguilar, Sergio Torres
contents This paper introduces TRIDIS (Tria Digita Scribunt), an open-source corpus of medieval and early modern manuscripts. TRIDIS aggregates multiple legacy collections (all published under open licenses) and incorporates large metadata descriptions. While prior publications referenced some portions of this corpus, here we provide a unified overview with a stronger focus on its constitution. We describe (i) the narrative, chronological, and editorial background of each major sub-corpus, (ii) its semi-diplomatic transcription rules (expansion, normalization, punctuation), (iii) a strategy for challenging out-of-domain test splits driven by outlier detection in a joint embedding space, and (iv) preliminary baseline experiments using TrOCR and MiniCPM2.5 comparing random and outlier-based test partitions. Overall, TRIDIS is designed to stimulate joint robust Handwritten Text Recognition (HTR) and Named Entity Recognition (NER) research across medieval and early modern textual heritage.
format Preprint
id arxiv_https___arxiv_org_abs_2503_22714
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TRIDIS: A Comprehensive Medieval and Early Modern Corpus for HTR and NER
Aguilar, Sergio Torres
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Digital Libraries
This paper introduces TRIDIS (Tria Digita Scribunt), an open-source corpus of medieval and early modern manuscripts. TRIDIS aggregates multiple legacy collections (all published under open licenses) and incorporates large metadata descriptions. While prior publications referenced some portions of this corpus, here we provide a unified overview with a stronger focus on its constitution. We describe (i) the narrative, chronological, and editorial background of each major sub-corpus, (ii) its semi-diplomatic transcription rules (expansion, normalization, punctuation), (iii) a strategy for challenging out-of-domain test splits driven by outlier detection in a joint embedding space, and (iv) preliminary baseline experiments using TrOCR and MiniCPM2.5 comparing random and outlier-based test partitions. Overall, TRIDIS is designed to stimulate joint robust Handwritten Text Recognition (HTR) and Named Entity Recognition (NER) research across medieval and early modern textual heritage.
title TRIDIS: A Comprehensive Medieval and Early Modern Corpus for HTR and NER
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Digital Libraries
url https://arxiv.org/abs/2503.22714