Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers
Fuente:
arXiv
Saved in:
| Main Authors: | Achara, Akshit, Chhabra, Anshuman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence Estimation
by: Vitel, Dmytro, et al.
Published: (2025)
by: Vitel, Dmytro, et al.
Published: (2025)
Revisiting Zero-Shot Abstractive Summarization in the Era of Large Language Models from the Perspective of Position Bias
by: Chhabra, Anshuman, et al.
Published: (2024)
by: Chhabra, Anshuman, et al.
Published: (2024)
Rethinking Reasoning in LLMs: Neuro-Symbolic Local RetoMaton Beyond ICL and CoT
by: Mamidala, Rushitha Santhoshi, et al.
Published: (2025)
by: Mamidala, Rushitha Santhoshi, et al.
Published: (2025)
Revealing the Underlying Patterns: Investigating Dataset Similarity, Performance, and Generalization
by: Achara, Akshit, et al.
Published: (2023)
by: Achara, Akshit, et al.
Published: (2023)
CoreDeep: Improving Crack Detection Algorithms Using Width Stochasticity
by: Pandey, Ram Krishna, et al.
Published: (2022)
by: Pandey, Ram Krishna, et al.
Published: (2022)
LLM Rationalis? Measuring Bargaining Capabilities of AI Negotiators
by: Shah, Cheril, et al.
Published: (2025)
by: Shah, Cheril, et al.
Published: (2025)
Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
by: Monjur, Ocean, et al.
Published: (2026)
by: Monjur, Ocean, et al.
Published: (2026)
Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models
by: Bonagiri, Akash, et al.
Published: (2025)
by: Bonagiri, Akash, et al.
Published: (2025)
Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
by: Nahin, Shahriar Kabir, et al.
Published: (2025)
by: Nahin, Shahriar Kabir, et al.
Published: (2025)
Debiasing Text Safety Classifiers through a Fairness-Aware Ensemble
by: Sturman, Olivia, et al.
Published: (2024)
by: Sturman, Olivia, et al.
Published: (2024)
LayerIF: Estimating Layer Quality for Large Language Models using Influence Functions
by: Askari, Hadi, et al.
Published: (2025)
by: Askari, Hadi, et al.
Published: (2025)
OMNIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and Modalities
by: Verma, Sahil, et al.
Published: (2025)
by: Verma, Sahil, et al.
Published: (2025)
Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation
by: Wróbel, Krzysztof, et al.
Published: (2026)
by: Wróbel, Krzysztof, et al.
Published: (2026)
Efficient Biomedical Entity Linking: Clinical Text Standardization with Low-Resource Techniques
by: Achara, Akshit, et al.
Published: (2024)
by: Achara, Akshit, et al.
Published: (2024)
Re-ranking Using Large Language Models for Mitigating Exposure to Harmful Content on Social Media Platforms
by: Oak, Rajvardhan, et al.
Published: (2025)
by: Oak, Rajvardhan, et al.
Published: (2025)
Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast
by: Avci, Mehmet Yigit, et al.
Published: (2026)
by: Avci, Mehmet Yigit, et al.
Published: (2026)
Classifying Human-Generated and AI-Generated Election Claims in Social Media
by: Dmonte, Alphaeus, et al.
Published: (2024)
by: Dmonte, Alphaeus, et al.
Published: (2024)
A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models
by: Feier, Andrei Marian, et al.
Published: (2026)
by: Feier, Andrei Marian, et al.
Published: (2026)
LionGuard: Building a Contextualized Moderation Classifier to Tackle Localized Unsafe Content
by: Foo, Jessica, et al.
Published: (2024)
by: Foo, Jessica, et al.
Published: (2024)
Incentivizing News Consumption on Social Media Platforms Using Large Language Models and Realistic Bot Accounts
by: Askari, Hadi, et al.
Published: (2024)
by: Askari, Hadi, et al.
Published: (2024)
Addressing Bias in LLMs: Strategies and Application to Fair AI-based Recruitment
by: Peña, Alejandro, et al.
Published: (2025)
by: Peña, Alejandro, et al.
Published: (2025)
The Multilingual Divide and Its Impact on Global AI Safety
by: Peppin, Aidan, et al.
Published: (2025)
by: Peppin, Aidan, et al.
Published: (2025)
AISafetyLab: A Comprehensive Framework for AI Safety Evaluation and Improvement
by: Zhang, Zhexin, et al.
Published: (2025)
by: Zhang, Zhexin, et al.
Published: (2025)
Invisible Attributes, Visible Biases: Exploring Demographic Shortcuts in MRI-based Alzheimer's Disease Classification
by: Achara, Akshit, et al.
Published: (2025)
by: Achara, Akshit, et al.
Published: (2025)
Effect of Gender Fair Job Description on Generative AI Images
by: Böckling, Finn, et al.
Published: (2025)
by: Böckling, Finn, et al.
Published: (2025)
SLM as Guardian: Pioneering AI Safety with Small Language Models
by: Kwon, Ohjoon, et al.
Published: (2024)
by: Kwon, Ohjoon, et al.
Published: (2024)
Building Trust: Foundations of Security, Safety and Transparency in AI
by: Sidhpurwala, Huzaifa, et al.
Published: (2024)
by: Sidhpurwala, Huzaifa, et al.
Published: (2024)
From Perceptions To Evidence: Detecting AI-Generated Content In Turkish News Media With A Fine-Tuned Bert Classifier
by: Ozdemir, Ozancan
Published: (2026)
by: Ozdemir, Ozancan
Published: (2026)
Introducing v0.5 of the AI Safety Benchmark from MLCommons
by: Vidgen, Bertie, et al.
Published: (2024)
by: Vidgen, Bertie, et al.
Published: (2024)
Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation
by: Zhang, Shutong, et al.
Published: (2026)
by: Zhang, Shutong, et al.
Published: (2026)
Are AI-Generated Text Detectors Robust to Adversarial Perturbations?
by: Huang, Guanhua, et al.
Published: (2024)
by: Huang, Guanhua, et al.
Published: (2024)
AI Knows When It's Being Watched: Functional Strategic Action and Contextual Register Modulation in Large Language Models
by: Covas, Vinicius, et al.
Published: (2026)
by: Covas, Vinicius, et al.
Published: (2026)
Watch Your Language: Investigating Content Moderation with Large Language Models
by: Kumar, Deepak, et al.
Published: (2023)
by: Kumar, Deepak, et al.
Published: (2023)
The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
by: Rios-Sialer, Ian
Published: (2026)
by: Rios-Sialer, Ian
Published: (2026)
AI Content Moderation in Therapy Conversations
by: Kim, Jiwon, et al.
Published: (2026)
by: Kim, Jiwon, et al.
Published: (2026)
Whose Preferences? Differences in Fairness Preferences and Their Impact on the Fairness of AI Utilizing Human Feedback
by: Lerner, Emilia Agis, et al.
Published: (2024)
by: Lerner, Emilia Agis, et al.
Published: (2024)
Doppelgänger's Watch: A Split Objective Approach to Large Language Models
by: Ghasemlou, Shervin, et al.
Published: (2024)
by: Ghasemlou, Shervin, et al.
Published: (2024)
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
Case-Aware LLM-as-a-Judge Evaluation for Enterprise-Scale RAG Systems
by: Chhabra, Mukul, et al.
Published: (2026)
by: Chhabra, Mukul, et al.
Published: (2026)
A Comparative Analysis of Counterfactual Explanation Methods for Text Classifiers
by: McAleese, Stephen, et al.
Published: (2024)
by: McAleese, Stephen, et al.
Published: (2024)
Similar Items
-
First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence Estimation
by: Vitel, Dmytro, et al.
Published: (2025) -
Revisiting Zero-Shot Abstractive Summarization in the Era of Large Language Models from the Perspective of Position Bias
by: Chhabra, Anshuman, et al.
Published: (2024) -
Rethinking Reasoning in LLMs: Neuro-Symbolic Local RetoMaton Beyond ICL and CoT
by: Mamidala, Rushitha Santhoshi, et al.
Published: (2025) -
Revealing the Underlying Patterns: Investigating Dataset Similarity, Performance, and Generalization
by: Achara, Akshit, et al.
Published: (2023) -
CoreDeep: Improving Crack Detection Algorithms Using Width Stochasticity
by: Pandey, Ram Krishna, et al.
Published: (2022)