The Enforcement and Feasibility of Hate Speech Moderation on Twitter

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tonneau, Manuel, Thurgood, Dylan, Liu, Diyi, Malhotra, Niyati, Orozco-Olvera, Victor, Schroeder, Ralph, Hale, Scott A., Ribeiro, Manoel Horta, Röttger, Paul, Fraiberger, Samuel P.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910128013312000
author Tonneau, Manuel
Thurgood, Dylan
Liu, Diyi
Malhotra, Niyati
Orozco-Olvera, Victor
Schroeder, Ralph
Hale, Scott A.
Ribeiro, Manoel Horta
Röttger, Paul
Fraiberger, Samuel P.
author_facet Tonneau, Manuel
Thurgood, Dylan
Liu, Diyi
Malhotra, Niyati
Orozco-Olvera, Victor
Schroeder, Ralph
Hale, Scott A.
Ribeiro, Manoel Horta
Röttger, Paul
Fraiberger, Samuel P.
contents Online hate speech is associated with substantial social harms, yet it remains unclear how consistently platforms enforce hate speech policies or whether enforcement is feasible at scale. We address these questions through a global audit of hate speech moderation on Twitter (now X). Using a complete 24-hour snapshot of public tweets, we construct representative samples comprising 540,000 tweets annotated for hate speech by trained annotators across eight major languages. Five months after posting, 80% of hateful tweets remain online, including explicitly violent hate speech. Such tweets are no more likely to be removed than non-hateful tweets, with neither severity nor visibility increasing the likelihood of removal. We then examine whether these enforcement gaps reflect technical limits of large-scale moderation systems. While fully automated detection systems cannot reliably identify hate speech without generating large numbers of false positives, they effectively prioritize likely violations for human review. Simulations of a human-AI moderation pipeline indicate that substantially reducing user exposure to hate speech is economically feasible at a cost below existing regulatory penalties. These results suggest that the persistence of online hate cannot be explained by technical constraints alone but also reflects institutional choices in the allocation of moderation resources.
format Preprint
id arxiv_https___arxiv_org_abs_2604_12289
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Enforcement and Feasibility of Hate Speech Moderation on Twitter
Tonneau, Manuel
Thurgood, Dylan
Liu, Diyi
Malhotra, Niyati
Orozco-Olvera, Victor
Schroeder, Ralph
Hale, Scott A.
Ribeiro, Manoel Horta
Röttger, Paul
Fraiberger, Samuel P.
Computers and Society
Computation and Language
Online hate speech is associated with substantial social harms, yet it remains unclear how consistently platforms enforce hate speech policies or whether enforcement is feasible at scale. We address these questions through a global audit of hate speech moderation on Twitter (now X). Using a complete 24-hour snapshot of public tweets, we construct representative samples comprising 540,000 tweets annotated for hate speech by trained annotators across eight major languages. Five months after posting, 80% of hateful tweets remain online, including explicitly violent hate speech. Such tweets are no more likely to be removed than non-hateful tweets, with neither severity nor visibility increasing the likelihood of removal. We then examine whether these enforcement gaps reflect technical limits of large-scale moderation systems. While fully automated detection systems cannot reliably identify hate speech without generating large numbers of false positives, they effectively prioritize likely violations for human review. Simulations of a human-AI moderation pipeline indicate that substantially reducing user exposure to hate speech is economically feasible at a cost below existing regulatory penalties. These results suggest that the persistence of online hate cannot be explained by technical constraints alone but also reflects institutional choices in the allocation of moderation resources.
title The Enforcement and Feasibility of Hate Speech Moderation on Twitter
topic Computers and Society
Computation and Language
url https://arxiv.org/abs/2604.12289