Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations
Fuente:
arXiv
Saved in:
| Main Authors: | Hartmann, David, Oueslati, Amin, Staufer, Dimitri, Pohlmann, Lena, Munzert, Simon, Heuer, Hendrik |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Watching the Watchers: A Comparative Fairness Audit of Cloud-based Content Moderation Services
by: Hartmann, David, et al.
Published: (2024)
by: Hartmann, David, et al.
Published: (2024)
HateModerate: Testing Hate Speech Detectors against Content Moderation Policies
by: Zheng, Jiangrui, et al.
Published: (2023)
by: Zheng, Jiangrui, et al.
Published: (2023)
Take the Power Back: Screen-Based Personal Moderation Against Hate Speech on Instagram
by: Luther, Anna Ricarda, et al.
Published: (2026)
by: Luther, Anna Ricarda, et al.
Published: (2026)
Beyond Hate: Differentiating Uncivil and Intolerant Speech in Multimodal Content Moderation
by: Herrmann, Nils A., et al.
Published: (2026)
by: Herrmann, Nils A., et al.
Published: (2026)
The Enforcement and Feasibility of Hate Speech Moderation on Twitter
by: Tonneau, Manuel, et al.
Published: (2026)
by: Tonneau, Manuel, et al.
Published: (2026)
HateBuffer: Safeguarding Content Moderators' Mental Well-Being through Hate Speech Content Modification
by: Park, Subin, et al.
Published: (2025)
by: Park, Subin, et al.
Published: (2025)
Silencing Empowerment, Allowing Bigotry: Auditing the Moderation of Hate Speech on Twitch
by: Shukla, Prarabdh, et al.
Published: (2025)
by: Shukla, Prarabdh, et al.
Published: (2025)
Recent Advances in Hate Speech Moderation: Multimodality and the Role of Large Models
by: Hee, Ming Shan, et al.
Published: (2024)
by: Hee, Ming Shan, et al.
Published: (2024)
Content Moderation Futures
by: Blackwell, Lindsay
Published: (2025)
by: Blackwell, Lindsay
Published: (2025)
Identity-related Speech Suppression in Generative AI Content Moderation
by: Proebsting, Grace, et al.
Published: (2024)
by: Proebsting, Grace, et al.
Published: (2024)
Explainability and Hate Speech: Structured Explanations Make Social Media Moderators Faster
by: Calabrese, Agostina, et al.
Published: (2024)
by: Calabrese, Agostina, et al.
Published: (2024)
Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions
by: Qu, Yiting, et al.
Published: (2025)
by: Qu, Yiting, et al.
Published: (2025)
AI Content Moderation in Therapy Conversations
by: Kim, Jiwon, et al.
Published: (2026)
by: Kim, Jiwon, et al.
Published: (2026)
NoisyHate: Mining Online Human-Written Perturbations for Realistic Robustness Benchmarking of Content Moderation Models
by: Ye, Yiran, et al.
Published: (2023)
by: Ye, Yiran, et al.
Published: (2023)
Ideology-Based LLMs for Content Moderation
by: Civelli, Stefano, et al.
Published: (2025)
by: Civelli, Stefano, et al.
Published: (2025)
Strategic Filtering for Content Moderation: Free Speech or Free of Distortion?
by: Ahmadi, Saba, et al.
Published: (2025)
by: Ahmadi, Saba, et al.
Published: (2025)
A Hate Speech Moderated Chat Application: Use Case for GDPR and DSA Compliance
by: Fillies, Jan, et al.
Published: (2024)
by: Fillies, Jan, et al.
Published: (2024)
Algorithmic Arbitrariness in Content Moderation
by: Gomez, Juan Felipe, et al.
Published: (2024)
by: Gomez, Juan Felipe, et al.
Published: (2024)
What Should LLMs Forget? Quantifying Personal Data in LLMs for Right-to-Be-Forgotten Requests
by: Staufer, Dimitri
Published: (2025)
by: Staufer, Dimitri
Published: (2025)
Longitudinal Monitoring of LLM Content Moderation of Social Issues
by: Dai, Yunlang, et al.
Published: (2025)
by: Dai, Yunlang, et al.
Published: (2025)
Multimodal Guidance Network for Missing-Modality Inference in Content Moderation
by: Zhao, Zhuokai, et al.
Published: (2023)
by: Zhao, Zhuokai, et al.
Published: (2023)
Personalizing Content Moderation on Social Media: User Perspectives on Moderation Choices, Interface Design, and Labor
by: Jhaver, Shagun, et al.
Published: (2023)
by: Jhaver, Shagun, et al.
Published: (2023)
Moderation Matters:Measuring Conversational Moderation Impact in English as a Second Language Group Discussion
by: Gao, Rena, et al.
Published: (2025)
by: Gao, Rena, et al.
Published: (2025)
Personalized Content Moderation and Emergent Outcomes
by: Gurkan, Necdet, et al.
Published: (2024)
by: Gurkan, Necdet, et al.
Published: (2024)
Improving Regulatory Oversight in Online Content Moderation
by: Tessa, Benedetta, et al.
Published: (2025)
by: Tessa, Benedetta, et al.
Published: (2025)
Wellbeing-Centered UX: Supporting Content Moderators
by: Mihalache, Diana, et al.
Published: (2025)
by: Mihalache, Diana, et al.
Published: (2025)
Content Moderation by LLM: From Accuracy to Legitimacy
by: Huang, Tao
Published: (2024)
by: Huang, Tao
Published: (2024)
Experimentation in Content Moderation using RWKV
by: Yildirim, Umut, et al.
Published: (2024)
by: Yildirim, Umut, et al.
Published: (2024)
Critical Challenges in Content Moderation for People Who Use Drugs (PWUD): Insights into Online Harm Reduction Practices from Moderators
by: Wang, Kaixuan, et al.
Published: (2025)
by: Wang, Kaixuan, et al.
Published: (2025)
"Ignorance is Not Bliss": Designing Personalized Moderation to Address Ableist Hate on Social Media
by: Heung, Sharon, et al.
Published: (2025)
by: Heung, Sharon, et al.
Published: (2025)
LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation
by: Park, Junyeong, et al.
Published: (2025)
by: Park, Junyeong, et al.
Published: (2025)
ExpGuard: LLM Content Moderation in Specialized Domains
by: Choi, Minseok, et al.
Published: (2026)
by: Choi, Minseok, et al.
Published: (2026)
Exploring the Boundaries of Content Moderation in Text-to-Image Generation
by: Riccio, Piera, et al.
Published: (2024)
by: Riccio, Piera, et al.
Published: (2024)
Probing Association Biases in LLM Moderation Over-Sensitivity
by: Wang, Yuxin, et al.
Published: (2025)
by: Wang, Yuxin, et al.
Published: (2025)
Human-Centred LLM Privacy Audits: Findings and Frictions
by: Staufer, Dimitri, et al.
Published: (2026)
by: Staufer, Dimitri, et al.
Published: (2026)
Moderating New Waves of Online Hate with Chain-of-Thought Reasoning in Large Language Models
by: Vishwamitra, Nishant, et al.
Published: (2023)
by: Vishwamitra, Nishant, et al.
Published: (2023)
What Do LLMs Associate with Your Name? A Human-Centered Black-Box Audit of Personal Data
by: Staufer, Dimitri, et al.
Published: (2026)
by: Staufer, Dimitri, et al.
Published: (2026)
Designing Child-Centered Content Exposure and Moderation
by: Saldías, Belén
Published: (2024)
by: Saldías, Belén
Published: (2024)
See, Explain, and Intervene: A Few-Shot Multimodal Agent Framework for Hateful Meme Moderation
by: Rizwan, Naquee, et al.
Published: (2026)
by: Rizwan, Naquee, et al.
Published: (2026)
Metamorphic Testing for Audio Content Moderation Software
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Similar Items
-
Watching the Watchers: A Comparative Fairness Audit of Cloud-based Content Moderation Services
by: Hartmann, David, et al.
Published: (2024) -
HateModerate: Testing Hate Speech Detectors against Content Moderation Policies
by: Zheng, Jiangrui, et al.
Published: (2023) -
Take the Power Back: Screen-Based Personal Moderation Against Hate Speech on Instagram
by: Luther, Anna Ricarda, et al.
Published: (2026) -
Beyond Hate: Differentiating Uncivil and Intolerant Speech in Multimodal Content Moderation
by: Herrmann, Nils A., et al.
Published: (2026) -
The Enforcement and Feasibility of Hate Speech Moderation on Twitter
by: Tonneau, Manuel, et al.
Published: (2026)