Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Quaremba, Gerrit, Rechkemmer, Amy, Black, Elizabeth, Vrandečić, Denny, Simperl, Elena
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918531427205120
author Quaremba, Gerrit
Rechkemmer, Amy
Black, Elizabeth
Vrandečić, Denny
Simperl, Elena
author_facet Quaremba, Gerrit
Rechkemmer, Amy
Black, Elizabeth
Vrandečić, Denny
Simperl, Elena
contents In automated fact-checking (AFC), check-worthiness detection identifies claims requiring verification based on domain-specific criteria. On Wikipedia, this task instantiates as Citation Needed Detection (CND), which flags claims lacking supporting citations. However, existing research has largely overlooked lower-resource languages, and recent AFC pipelines rely on large language models (LLMs), which are inaccessible to low-resource organizations. We introduce MCN, a multilingual CND corpus spanning 18 languages across three resource levels, on which we conduct an extensive study of small decoder-based language models (SLMs). Our experiments show that SLMs fine-tuned with an encoder-style objective substantially outperform prompted LLMs across languages. We further present one of the first studies on cross-lingual CND, demonstrating that SLMs fine-tuned solely on English claims surpass LLMs, even with little to no target-language adaptation. Our findings have important implications for lower-resource Wikipedia communities and suggest that compact, task-specific models are preferable to LLMs for CND. We release all data and code at https://github.com/gerritq/mcn
format Preprint
id arxiv_https___arxiv_org_abs_2605_31136
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages
Quaremba, Gerrit
Rechkemmer, Amy
Black, Elizabeth
Vrandečić, Denny
Simperl, Elena
Computation and Language
In automated fact-checking (AFC), check-worthiness detection identifies claims requiring verification based on domain-specific criteria. On Wikipedia, this task instantiates as Citation Needed Detection (CND), which flags claims lacking supporting citations. However, existing research has largely overlooked lower-resource languages, and recent AFC pipelines rely on large language models (LLMs), which are inaccessible to low-resource organizations. We introduce MCN, a multilingual CND corpus spanning 18 languages across three resource levels, on which we conduct an extensive study of small decoder-based language models (SLMs). Our experiments show that SLMs fine-tuned with an encoder-style objective substantially outperform prompted LLMs across languages. We further present one of the first studies on cross-lingual CND, demonstrating that SLMs fine-tuned solely on English claims surpass LLMs, even with little to no target-language adaptation. Our findings have important implications for lower-resource Wikipedia communities and suggest that compact, task-specific models are preferable to LLMs for CND. We release all data and code at https://github.com/gerritq/mcn
title Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages
topic Computation and Language
url https://arxiv.org/abs/2605.31136