Watching the Watchers: A Comparative Fairness Audit of Cloud-based Content Moderation Services
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hartmann, David, Oueslati, Amin, Staufer, Dimitri |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations
von: Hartmann, David, et al.
Veröffentlicht: (2025)
von: Hartmann, David, et al.
Veröffentlicht: (2025)
Human-Centred LLM Privacy Audits: Findings and Frictions
von: Staufer, Dimitri, et al.
Veröffentlicht: (2026)
von: Staufer, Dimitri, et al.
Veröffentlicht: (2026)
What Do LLMs Associate with Your Name? A Human-Centered Black-Box Audit of Personal Data
von: Staufer, Dimitri, et al.
Veröffentlicht: (2026)
von: Staufer, Dimitri, et al.
Veröffentlicht: (2026)
The "Colonial Impulse" of Natural Language Processing: An Audit of Bengali Sentiment Analysis Tools and Their Identity-based Biases
von: Das, Dipto, et al.
Veröffentlicht: (2024)
von: Das, Dipto, et al.
Veröffentlicht: (2024)
Asking For It: Question-Answering for Predicting Rule Infractions in Online Content Moderation
von: Samory, Mattia, et al.
Veröffentlicht: (2025)
von: Samory, Mattia, et al.
Veröffentlicht: (2025)
Does Writing with Language Models Reduce Content Diversity?
von: Padmakumar, Vishakh, et al.
Veröffentlicht: (2023)
von: Padmakumar, Vishakh, et al.
Veröffentlicht: (2023)
Silencing Empowerment, Allowing Bigotry: Auditing the Moderation of Hate Speech on Twitch
von: Shukla, Prarabdh, et al.
Veröffentlicht: (2025)
von: Shukla, Prarabdh, et al.
Veröffentlicht: (2025)
Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce
von: Shao, Yijia, et al.
Veröffentlicht: (2025)
von: Shao, Yijia, et al.
Veröffentlicht: (2025)
Longitudinal Monitoring of LLM Content Moderation of Social Issues
von: Dai, Yunlang, et al.
Veröffentlicht: (2025)
von: Dai, Yunlang, et al.
Veröffentlicht: (2025)
Identity-related Speech Suppression in Generative AI Content Moderation
von: Proebsting, Grace, et al.
Veröffentlicht: (2024)
von: Proebsting, Grace, et al.
Veröffentlicht: (2024)
LLM-based Cognitive Models of Students with Misconceptions
von: Sonkar, Shashank, et al.
Veröffentlicht: (2024)
von: Sonkar, Shashank, et al.
Veröffentlicht: (2024)
The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
Is it Still Fair? A Comparative Evaluation of Fairness Algorithms through the Lens of Covariate Drift
von: Deho, Oscar Blessed, et al.
Veröffentlicht: (2024)
von: Deho, Oscar Blessed, et al.
Veröffentlicht: (2024)
Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications
von: Dong, Wenhan, et al.
Veröffentlicht: (2025)
von: Dong, Wenhan, et al.
Veröffentlicht: (2025)
Toward Cultural Interpretability: A Linguistic Anthropological Framework for Describing and Evaluating Large Language Models (LLMs)
von: Jones, Graham M., et al.
Veröffentlicht: (2024)
von: Jones, Graham M., et al.
Veröffentlicht: (2024)
Auditing African Content Moderators' Working Conditions by Using the European General Data Protection Regulation (GDPR)
von: Tighanimine, Mariame, et al.
Veröffentlicht: (2026)
von: Tighanimine, Mariame, et al.
Veröffentlicht: (2026)
Content Moderation Justice and Fairness on Social Media: Comparisons Across Different Contexts and Platforms
von: Cai, Jie, et al.
Veröffentlicht: (2024)
von: Cai, Jie, et al.
Veröffentlicht: (2024)
The Moral Gap of Large Language Models
von: Skorski, Maciej, et al.
Veröffentlicht: (2025)
von: Skorski, Maciej, et al.
Veröffentlicht: (2025)
Attention to Non-Adopters
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2025)
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2025)
Activation Steering via Generative Causal Mediation
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
Content Moderation by LLM: From Accuracy to Legitimacy
von: Huang, Tao
Veröffentlicht: (2024)
von: Huang, Tao
Veröffentlicht: (2024)
Content Moderation Futures
von: Blackwell, Lindsay
Veröffentlicht: (2025)
von: Blackwell, Lindsay
Veröffentlicht: (2025)
Fairness-Aware Few-Shot Learning for Audio-Visual Stress Detection
von: Shelke, Anushka Sanjay, et al.
Veröffentlicht: (2025)
von: Shelke, Anushka Sanjay, et al.
Veröffentlicht: (2025)
An Explainable and Fair AI Tool for PCOS Risk Assessment: Calibration, Subgroup Equity, and Interactive Clinical Deployment
von: Khan, Asma Sadia, et al.
Veröffentlicht: (2025)
von: Khan, Asma Sadia, et al.
Veröffentlicht: (2025)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
von: Zheng, Mingqian, et al.
Veröffentlicht: (2023)
von: Zheng, Mingqian, et al.
Veröffentlicht: (2023)
Sociodemographic Prompting is Not Yet an Effective Approach for Simulating Subjective Judgments with LLMs
von: Sun, Huaman, et al.
Veröffentlicht: (2023)
von: Sun, Huaman, et al.
Veröffentlicht: (2023)
The Impossibility of Fair LLMs
von: Anthis, Jacy, et al.
Veröffentlicht: (2024)
von: Anthis, Jacy, et al.
Veröffentlicht: (2024)
What Should LLMs Forget? Quantifying Personal Data in LLMs for Right-to-Be-Forgotten Requests
von: Staufer, Dimitri
Veröffentlicht: (2025)
von: Staufer, Dimitri
Veröffentlicht: (2025)
Implicit Personalization in Language Models: A Systematic Study
von: Jin, Zhijing, et al.
Veröffentlicht: (2024)
von: Jin, Zhijing, et al.
Veröffentlicht: (2024)
Representation Bias of Adolescents in AI: A Bilingual, Bicultural Study
von: Wolfe, Robert, et al.
Veröffentlicht: (2024)
von: Wolfe, Robert, et al.
Veröffentlicht: (2024)
AI Content Moderation in Therapy Conversations
von: Kim, Jiwon, et al.
Veröffentlicht: (2026)
von: Kim, Jiwon, et al.
Veröffentlicht: (2026)
MentalChat16K: A Benchmark Dataset for Conversational Mental Health Assistance
von: Xu, Jia, et al.
Veröffentlicht: (2025)
von: Xu, Jia, et al.
Veröffentlicht: (2025)
PsychoGAT: A Novel Psychological Measurement Paradigm through Interactive Fiction Games with LLM Agents
von: Yang, Qisen, et al.
Veröffentlicht: (2024)
von: Yang, Qisen, et al.
Veröffentlicht: (2024)
LLM4PM: A case study on using Large Language Models for Process Modeling in Enterprise Organizations
von: Ziche, Clara, et al.
Veröffentlicht: (2024)
von: Ziche, Clara, et al.
Veröffentlicht: (2024)
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate
von: Ngueajio, Mikel K., et al.
Veröffentlicht: (2025)
von: Ngueajio, Mikel K., et al.
Veröffentlicht: (2025)
Overreliance on AI in Information-seeking from Video Content
von: Møller, Anders Giovanni, et al.
Veröffentlicht: (2026)
von: Møller, Anders Giovanni, et al.
Veröffentlicht: (2026)
Are Open-Weight LLMs Ready for Social Media Moderation? A Comparative Study on Bluesky
von: Chou, Hsuan-Yu, et al.
Veröffentlicht: (2026)
von: Chou, Hsuan-Yu, et al.
Veröffentlicht: (2026)
Fairness-in-the-Workflow: How Machine Learning Practitioners at Big Tech Companies Approach Fairness in Recommender Systems
von: Yan, Jing Nathan, et al.
Veröffentlicht: (2025)
von: Yan, Jing Nathan, et al.
Veröffentlicht: (2025)
Effort-aware Fairness: Incorporating a Philosophy-informed, Human-centered Notion of Effort into Algorithmic Fairness Metrics
von: Nguyen, Tin Trung, et al.
Veröffentlicht: (2025)
von: Nguyen, Tin Trung, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations
von: Hartmann, David, et al.
Veröffentlicht: (2025) -
Human-Centred LLM Privacy Audits: Findings and Frictions
von: Staufer, Dimitri, et al.
Veröffentlicht: (2026) -
What Do LLMs Associate with Your Name? A Human-Centered Black-Box Audit of Personal Data
von: Staufer, Dimitri, et al.
Veröffentlicht: (2026) -
The "Colonial Impulse" of Natural Language Processing: An Audit of Bengali Sentiment Analysis Tools and Their Identity-based Biases
von: Das, Dipto, et al.
Veröffentlicht: (2024) -
Asking For It: Question-Answering for Predicting Rule Infractions in Online Content Moderation
von: Samory, Mattia, et al.
Veröffentlicht: (2025)