Language-Agnostic Modeling of Source Reliability on Wikipedia

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: D'Ignazi, Jacopo, Kaltenbrunner, Andreas, Mejova, Yelena, Tizzani, Michele, Kalimeri, Kyriaki, Beiró, Mariano, Aragón, Pablo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912719771271168
author D'Ignazi, Jacopo
Kaltenbrunner, Andreas
Mejova, Yelena
Tizzani, Michele
Kalimeri, Kyriaki
Beiró, Mariano
Aragón, Pablo
author_facet D'Ignazi, Jacopo
Kaltenbrunner, Andreas
Mejova, Yelena
Tizzani, Michele
Kalimeri, Kyriaki
Beiró, Mariano
Aragón, Pablo
contents Over the last few years, verifying the credibility of information sources has become a fundamental need to combat disinformation. Here, we present a language-agnostic model designed to assess the reliability of web domains as sources in references across multiple language editions of Wikipedia. Utilizing editing activity data, the model evaluates domain reliability within different articles of varying controversiality, such as Climate Change, COVID-19, History, Media, and Biology topics. Crafting features that express domain usage across articles, the model effectively predicts domain reliability, achieving an F1 Macro score of approximately 0.80 for English and other high-resource languages. For mid-resource languages, we achieve 0.65, while the performance of low-resource languages varies. In all cases, the time the domain remains present in the articles (which we dub as permanence) is one of the most predictive features. We highlight the challenge of maintaining consistent model performance across languages of varying resource levels and demonstrate that adapting models from higher-resource languages can improve performance. We believe these findings can assist Wikipedia editors in their ongoing efforts to verify citations and may offer useful insights for other user-generated content communities.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18803
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Language-Agnostic Modeling of Source Reliability on Wikipedia
D'Ignazi, Jacopo
Kaltenbrunner, Andreas
Mejova, Yelena
Tizzani, Michele
Kalimeri, Kyriaki
Beiró, Mariano
Aragón, Pablo
Social and Information Networks
Machine Learning
Over the last few years, verifying the credibility of information sources has become a fundamental need to combat disinformation. Here, we present a language-agnostic model designed to assess the reliability of web domains as sources in references across multiple language editions of Wikipedia. Utilizing editing activity data, the model evaluates domain reliability within different articles of varying controversiality, such as Climate Change, COVID-19, History, Media, and Biology topics. Crafting features that express domain usage across articles, the model effectively predicts domain reliability, achieving an F1 Macro score of approximately 0.80 for English and other high-resource languages. For mid-resource languages, we achieve 0.65, while the performance of low-resource languages varies. In all cases, the time the domain remains present in the articles (which we dub as permanence) is one of the most predictive features. We highlight the challenge of maintaining consistent model performance across languages of varying resource levels and demonstrate that adapting models from higher-resource languages can improve performance. We believe these findings can assist Wikipedia editors in their ongoing efforts to verify citations and may offer useful insights for other user-generated content communities.
title Language-Agnostic Modeling of Source Reliability on Wikipedia
topic Social and Information Networks
Machine Learning
url https://arxiv.org/abs/2410.18803