Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Achara, Akshit, Chhabra, Anshuman
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917900329156608
author Achara, Akshit
Chhabra, Anshuman
author_facet Achara, Akshit
Chhabra, Anshuman
contents AI Safety Moderation (ASM) classifiers are designed to moderate content on social media platforms and to serve as guardrails that prevent Large Language Models (LLMs) from being fine-tuned on unsafe inputs. Owing to their potential for disparate impact, it is crucial to ensure that these classifiers: (1) do not unfairly classify content belonging to users from minority groups as unsafe compared to those from majority groups and (2) that their behavior remains robust and consistent across similar inputs. In this work, we thus examine the fairness and robustness of four widely-used, closed-source ASM classifiers: OpenAI Moderation API, Perspective API, Google Cloud Natural Language (GCNL) API, and Clarifai API. We assess fairness using metrics such as demographic parity and conditional statistical parity, comparing their performance against ASM models and a fair-only baseline. Additionally, we analyze robustness by testing the classifiers' sensitivity to small and natural input perturbations. Our findings reveal potential fairness and robustness gaps, highlighting the need to mitigate these issues in future versions of these models.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13302
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers
Achara, Akshit
Chhabra, Anshuman
Computation and Language
Artificial Intelligence
AI Safety Moderation (ASM) classifiers are designed to moderate content on social media platforms and to serve as guardrails that prevent Large Language Models (LLMs) from being fine-tuned on unsafe inputs. Owing to their potential for disparate impact, it is crucial to ensure that these classifiers: (1) do not unfairly classify content belonging to users from minority groups as unsafe compared to those from majority groups and (2) that their behavior remains robust and consistent across similar inputs. In this work, we thus examine the fairness and robustness of four widely-used, closed-source ASM classifiers: OpenAI Moderation API, Perspective API, Google Cloud Natural Language (GCNL) API, and Clarifai API. We assess fairness using metrics such as demographic parity and conditional statistical parity, comparing their performance against ASM models and a fair-only baseline. Additionally, we analyze robustness by testing the classifiers' sensitivity to small and natural input perturbations. Our findings reveal potential fairness and robustness gaps, highlighting the need to mitigate these issues in future versions of these models.
title Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2501.13302