Quid est VERITAS? A Modular Framework for Archival Document Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bassanini, Leonardo, Biancardi, Ludovico, Ferrara, Alfio, Gamberini, Andrea, Picascia, Sergio, Vaglienti, Folco
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910084774232064
author Bassanini, Leonardo
Biancardi, Ludovico
Ferrara, Alfio
Gamberini, Andrea
Picascia, Sergio
Vaglienti, Folco
author_facet Bassanini, Leonardo
Biancardi, Ludovico
Ferrara, Alfio
Gamberini, Andrea
Picascia, Sergio
Vaglienti, Folco
contents The digitisation of historical documents has traditionally been conceived as a process limited to character-level transcription, producing flat text that lacks the structural and semantic information necessary for substantive computational analysis. We present VERITAS (Vision-Enhanced Reading, Interpretation, and Transcription of Archival Sources), a modular, model-agnostic framework that reconceptualises digitisation as an integrated workflow encompassing transcription, layout analysis, and semantic enrichment. The pipeline is organised into four stages - Preprocessing, Extraction, Refinement, and Enrichment - and employs a schema-driven architecture that allows researchers to declaratively specify their extraction objectives. We evaluate VERITAS on the critical edition of Bernardino Corio's Storia di Milano, a Renaissance chronicle of over 1,600 pages. Results demonstrate that the pipeline achieves a 67.6% relative reduction in word error rate compared to a commercial OCR baseline, with a threefold reduction in end-to-end processing time when accounting for manual correction. We further illustrate the downstream utility of the pipeline's output by querying the transcribed corpus through a retrieval-augmented generation system, demonstrating its capacity to support historical inquiry.
format Preprint
id arxiv_https___arxiv_org_abs_2603_28108
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Quid est VERITAS? A Modular Framework for Archival Document Analysis
Bassanini, Leonardo
Biancardi, Ludovico
Ferrara, Alfio
Gamberini, Andrea
Picascia, Sergio
Vaglienti, Folco
Digital Libraries
Artificial Intelligence
Information Retrieval
The digitisation of historical documents has traditionally been conceived as a process limited to character-level transcription, producing flat text that lacks the structural and semantic information necessary for substantive computational analysis. We present VERITAS (Vision-Enhanced Reading, Interpretation, and Transcription of Archival Sources), a modular, model-agnostic framework that reconceptualises digitisation as an integrated workflow encompassing transcription, layout analysis, and semantic enrichment. The pipeline is organised into four stages - Preprocessing, Extraction, Refinement, and Enrichment - and employs a schema-driven architecture that allows researchers to declaratively specify their extraction objectives. We evaluate VERITAS on the critical edition of Bernardino Corio's Storia di Milano, a Renaissance chronicle of over 1,600 pages. Results demonstrate that the pipeline achieves a 67.6% relative reduction in word error rate compared to a commercial OCR baseline, with a threefold reduction in end-to-end processing time when accounting for manual correction. We further illustrate the downstream utility of the pipeline's output by querying the transcribed corpus through a retrieval-augmented generation system, demonstrating its capacity to support historical inquiry.
title Quid est VERITAS? A Modular Framework for Archival Document Analysis
topic Digital Libraries
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2603.28108