Fetch-A-Set: A Large-Scale OCR-Free Benchmark for Historical Document Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Molina, Adrià, Terrades, Oriol Ramos, Lladós, Josep
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929387829460992
author Molina, Adrià
Terrades, Oriol Ramos
Lladós, Josep
author_facet Molina, Adrià
Terrades, Oriol Ramos
Lladós, Josep
contents This paper introduces Fetch-A-Set (FAS), a comprehensive benchmark tailored for legislative historical document analysis systems, addressing the challenges of large-scale document retrieval in historical contexts. The benchmark comprises a vast repository of documents dating back to the XVII century, serving both as a training resource and an evaluation benchmark for retrieval systems. It fills a critical gap in the literature by focusing on complex extractive tasks within the domain of cultural heritage. The proposed benchmark tackles the multifaceted problem of historical document analysis, including text-to-image retrieval for queries and image-to-text topic extraction from document fragments, all while accommodating varying levels of document legibility. This benchmark aims to spur advancements in the field by providing baselines and data for the development and evaluation of robust historical document retrieval systems, particularly in scenarios characterized by wide historical spectrum.
format Preprint
id arxiv_https___arxiv_org_abs_2406_07315
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fetch-A-Set: A Large-Scale OCR-Free Benchmark for Historical Document Retrieval
Molina, Adrià
Terrades, Oriol Ramos
Lladós, Josep
Information Retrieval
Computer Vision and Pattern Recognition
This paper introduces Fetch-A-Set (FAS), a comprehensive benchmark tailored for legislative historical document analysis systems, addressing the challenges of large-scale document retrieval in historical contexts. The benchmark comprises a vast repository of documents dating back to the XVII century, serving both as a training resource and an evaluation benchmark for retrieval systems. It fills a critical gap in the literature by focusing on complex extractive tasks within the domain of cultural heritage. The proposed benchmark tackles the multifaceted problem of historical document analysis, including text-to-image retrieval for queries and image-to-text topic extraction from document fragments, all while accommodating varying levels of document legibility. This benchmark aims to spur advancements in the field by providing baselines and data for the development and evaluation of robust historical document retrieval systems, particularly in scenarios characterized by wide historical spectrum.
title Fetch-A-Set: A Large-Scale OCR-Free Benchmark for Historical Document Retrieval
topic Information Retrieval
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.07315