What Large Language Models Do Not Talk About: An Empirical Study of Moderation and Censorship Practices

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Noels, Sander, Bied, Guillaume, Buyl, Maarten, Rogiers, Alexander, Fettach, Yousra, Lijffijt, Jefrey, De Bie, Tijl
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908302687862784
author Noels, Sander
Bied, Guillaume
Buyl, Maarten
Rogiers, Alexander
Fettach, Yousra
Lijffijt, Jefrey
De Bie, Tijl
author_facet Noels, Sander
Bied, Guillaume
Buyl, Maarten
Rogiers, Alexander
Fettach, Yousra
Lijffijt, Jefrey
De Bie, Tijl
contents Large Language Models (LLMs) are increasingly deployed as gateways to information, yet their content moderation practices remain underexplored. This work investigates the extent to which LLMs refuse to answer or omit information when prompted on political topics. To do so, we distinguish between hard censorship (i.e., generated refusals, error messages, or canned denial responses) and soft censorship (i.e., selective omission or downplaying of key elements), which we identify in LLMs' responses when asked to provide information on a broad range of political figures. Our analysis covers 14 state-of-the-art models from Western countries, China, and Russia, prompted in all six official United Nations (UN) languages. Our analysis suggests that although censorship is observed across the board, it is predominantly tailored to an LLM provider's domestic audience and typically manifests as either hard censorship or soft censorship (though rarely both concurrently). These findings underscore the need for ideological and geographic diversity among publicly available LLMs, and greater transparency in LLM moderation strategies to facilitate informed user choices. All data are made freely available.
format Preprint
id arxiv_https___arxiv_org_abs_2504_03803
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What Large Language Models Do Not Talk About: An Empirical Study of Moderation and Censorship Practices
Noels, Sander
Bied, Guillaume
Buyl, Maarten
Rogiers, Alexander
Fettach, Yousra
Lijffijt, Jefrey
De Bie, Tijl
Computation and Language
Computers and Society
Machine Learning
Large Language Models (LLMs) are increasingly deployed as gateways to information, yet their content moderation practices remain underexplored. This work investigates the extent to which LLMs refuse to answer or omit information when prompted on political topics. To do so, we distinguish between hard censorship (i.e., generated refusals, error messages, or canned denial responses) and soft censorship (i.e., selective omission or downplaying of key elements), which we identify in LLMs' responses when asked to provide information on a broad range of political figures. Our analysis covers 14 state-of-the-art models from Western countries, China, and Russia, prompted in all six official United Nations (UN) languages. Our analysis suggests that although censorship is observed across the board, it is predominantly tailored to an LLM provider's domestic audience and typically manifests as either hard censorship or soft censorship (though rarely both concurrently). These findings underscore the need for ideological and geographic diversity among publicly available LLMs, and greater transparency in LLM moderation strategies to facilitate informed user choices. All data are made freely available.
title What Large Language Models Do Not Talk About: An Empirical Study of Moderation and Censorship Practices
topic Computation and Language
Computers and Society
Machine Learning
url https://arxiv.org/abs/2504.03803