Saved in:
Bibliographic Details
Main Authors: De La Fuente-Cuesta, Alejandro, Martinez-Serra, Alberto, Visscher, Nienke, Castro, Laia, Cardenal, Ana S.
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.17435
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912687137488896
author De La Fuente-Cuesta, Alejandro
Martinez-Serra, Alberto
Visscher, Nienke
Castro, Laia
Cardenal, Ana S.
author_facet De La Fuente-Cuesta, Alejandro
Martinez-Serra, Alberto
Visscher, Nienke
Castro, Laia
Cardenal, Ana S.
contents The use of large language models (LLMs) is becoming common in political science and digital media research. While LLMs have demonstrated ability in labelling tasks, their effectiveness to classify Political Content (PC) from URLs remains underexplored. This article evaluates whether LLMs can accurately distinguish PC from non-PC using both the text and the URLs of news articles across five countries (France, Germany, Spain, the UK, and the US) and their different languages. Using cutting-edge models, we benchmark their performance against human-coded data to assess whether URL-level analysis can approximate full-text analysis. Our findings show that URLs embed relevant information and can serve as a scalable, cost-effective alternative to discern PC. However, we also uncover systematic biases: LLMs seem to overclassify centrist news as political, leading to false positives that may distort further analyses. We conclude by outlining methodological recommendations on the use of LLMs in political science research.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17435
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond the Link: Assessing LLMs' ability to Classify Political Content across Global Media
De La Fuente-Cuesta, Alejandro
Martinez-Serra, Alberto
Visscher, Nienke
Castro, Laia
Cardenal, Ana S.
Computation and Language
The use of large language models (LLMs) is becoming common in political science and digital media research. While LLMs have demonstrated ability in labelling tasks, their effectiveness to classify Political Content (PC) from URLs remains underexplored. This article evaluates whether LLMs can accurately distinguish PC from non-PC using both the text and the URLs of news articles across five countries (France, Germany, Spain, the UK, and the US) and their different languages. Using cutting-edge models, we benchmark their performance against human-coded data to assess whether URL-level analysis can approximate full-text analysis. Our findings show that URLs embed relevant information and can serve as a scalable, cost-effective alternative to discern PC. However, we also uncover systematic biases: LLMs seem to overclassify centrist news as political, leading to false positives that may distort further analyses. We conclude by outlining methodological recommendations on the use of LLMs in political science research.
title Beyond the Link: Assessing LLMs' ability to Classify Political Content across Global Media
topic Computation and Language
url https://arxiv.org/abs/2506.17435