Leveraging Digitized Newspapers to Collect Summarization Data in Low-Resource Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dahan, Noam, Kidron, Omer, Stanovsky, Gabriel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911274412015616
author Dahan, Noam
Kidron, Omer
Stanovsky, Gabriel
author_facet Dahan, Noam
Kidron, Omer
Stanovsky, Gabriel
contents High quality summarization data remains scarce in under-represented languages. However, historical newspapers, made available through recent digitization efforts, offer an abundant source of untapped, naturally annotated data. In this work, we present a novel method for collecting naturally occurring summaries via Front-Page Teasers, where editors summarize full length articles. We show that this phenomenon is common across seven diverse languages and supports multi-document summarization. To scale data collection, we develop an automatic process, suited to varying linguistic resource levels. Finally, we apply this process to a Hebrew newspaper title, producing HEBTEASESUM, the first dedicated multi-document summarization dataset in Hebrew.
format Preprint
id arxiv_https___arxiv_org_abs_2511_14598
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Digitized Newspapers to Collect Summarization Data in Low-Resource Languages
Dahan, Noam
Kidron, Omer
Stanovsky, Gabriel
Computation and Language
High quality summarization data remains scarce in under-represented languages. However, historical newspapers, made available through recent digitization efforts, offer an abundant source of untapped, naturally annotated data. In this work, we present a novel method for collecting naturally occurring summaries via Front-Page Teasers, where editors summarize full length articles. We show that this phenomenon is common across seven diverse languages and supports multi-document summarization. To scale data collection, we develop an automatic process, suited to varying linguistic resource levels. Finally, we apply this process to a Hebrew newspaper title, producing HEBTEASESUM, the first dedicated multi-document summarization dataset in Hebrew.
title Leveraging Digitized Newspapers to Collect Summarization Data in Low-Resource Languages
topic Computation and Language
url https://arxiv.org/abs/2511.14598