Towards Safer AI Moderation: Evaluating LLM Moderators Through a Unified Benchmark Dataset and Advocating a Human-First Approach
Fuente:
arXiv
Salvato in:
| Autori principali: | Machlovi, Naseem, Saleki, Maryam, Ababio, Innocent, Amin, Ruhul |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Multi-Perspective Benchmark and Moderation Model for Evaluating Safety and Adversarial Robustness
di: Machlovi, Naseem, et al.
Pubblicazione: (2025)
di: Machlovi, Naseem, et al.
Pubblicazione: (2025)
Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models
di: Ren, Juan, et al.
Pubblicazione: (2025)
di: Ren, Juan, et al.
Pubblicazione: (2025)
Redefining Elderly Care with Agentic AI: Challenges and Opportunities
di: Khalil, Ruhul Amin, et al.
Pubblicazione: (2025)
di: Khalil, Ruhul Amin, et al.
Pubblicazione: (2025)
VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare
di: Shetty, Anudeex, et al.
Pubblicazione: (2025)
di: Shetty, Anudeex, et al.
Pubblicazione: (2025)
AI Benchmarks and Datasets for LLM Evaluation
di: Ivanov, Todor, et al.
Pubblicazione: (2024)
di: Ivanov, Todor, et al.
Pubblicazione: (2024)
AI & Human Co-Improvement for Safer Co-Superintelligence
di: Weston, Jason, et al.
Pubblicazione: (2025)
di: Weston, Jason, et al.
Pubblicazione: (2025)
Towards a Humanized Social-Media Ecosystem: AI-Augmented HCI Design Patterns for Safety, Agency & Well-Being
di: Ameen, Mohd Ruhul, et al.
Pubblicazione: (2025)
di: Ameen, Mohd Ruhul, et al.
Pubblicazione: (2025)
Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
di: Kumar, Shanu, et al.
Pubblicazione: (2024)
di: Kumar, Shanu, et al.
Pubblicazione: (2024)
Evaluating Online Moderation Via LLM-Powered Counterfactual Simulations
di: Fidone, Giacomo, et al.
Pubblicazione: (2025)
di: Fidone, Giacomo, et al.
Pubblicazione: (2025)
Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models
di: Bonagiri, Akash, et al.
Pubblicazione: (2025)
di: Bonagiri, Akash, et al.
Pubblicazione: (2025)
Toward Trustworthy Evaluation of Sustainability Rating Methodologies: A Human-AI Collaborative Framework for Benchmark Dataset Construction
di: Cai, Xiaoran, et al.
Pubblicazione: (2026)
di: Cai, Xiaoran, et al.
Pubblicazione: (2026)
Advancing Embodied Agent Security: From Safety Benchmarks to Input Moderation
di: Wang, Ning, et al.
Pubblicazione: (2025)
di: Wang, Ning, et al.
Pubblicazione: (2025)
Devil's Advocate: Anticipatory Reflection for LLM Agents
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
Probing Association Biases in LLM Moderation Over-Sensitivity
di: Wang, Yuxin, et al.
Pubblicazione: (2025)
di: Wang, Yuxin, et al.
Pubblicazione: (2025)
LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation
di: Park, Junyeong, et al.
Pubblicazione: (2025)
di: Park, Junyeong, et al.
Pubblicazione: (2025)
Tricky$^2$: Towards a Benchmark for Evaluating Human and LLM Error Interactions
di: Granger, Cole, et al.
Pubblicazione: (2026)
di: Granger, Cole, et al.
Pubblicazione: (2026)
FLAME: Flexible LLM-Assisted Moderation Engine
di: Bakulin, Ivan, et al.
Pubblicazione: (2025)
di: Bakulin, Ivan, et al.
Pubblicazione: (2025)
Competing LLM Agents in a Non-Cooperative Game of Opinion Polarisation
di: Qasmi, Amin, et al.
Pubblicazione: (2025)
di: Qasmi, Amin, et al.
Pubblicazione: (2025)
Detecting AI-Generated Images via Diffusion Snap-Back Reconstruction: A Forensic Approach
di: Ameen, Mohd Ruhul, et al.
Pubblicazione: (2025)
di: Ameen, Mohd Ruhul, et al.
Pubblicazione: (2025)
LLM Performance Predictors: Learning When to Escalate in Hybrid Human-AI Moderation Systems
di: Bachar, Or, et al.
Pubblicazione: (2026)
di: Bachar, Or, et al.
Pubblicazione: (2026)
Metagoals Endowing Self-Modifying AGI Systems with Goal Stability or Moderated Goal Evolution: Toward a Formally Sound and Practical Approach
di: Goertzel, Ben
Pubblicazione: (2024)
di: Goertzel, Ben
Pubblicazione: (2024)
GMP: A Benchmark for Content Moderation under Co-occurring Violations and Dynamic Rules
di: Dong, Houde, et al.
Pubblicazione: (2026)
di: Dong, Houde, et al.
Pubblicazione: (2026)
AI as Equalizer or Amplifier? Task Complexity as the Moderating Factor for Human Expertise in Hybrid Intelligence Systems
di: An, Tao
Pubblicazione: (2025)
di: An, Tao
Pubblicazione: (2025)
OMNIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and Modalities
di: Verma, Sahil, et al.
Pubblicazione: (2025)
di: Verma, Sahil, et al.
Pubblicazione: (2025)
Lessons Learned from Evaluation of LLM based Multi-agents in Safer Therapy Recommendation
di: Wu, Yicong, et al.
Pubblicazione: (2025)
di: Wu, Yicong, et al.
Pubblicazione: (2025)
Towards Competent AI for Fundamental Analysis in Finance: A Benchmark Dataset and Evaluation
di: Wu, Zonghan, et al.
Pubblicazione: (2025)
di: Wu, Zonghan, et al.
Pubblicazione: (2025)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
di: Atasoy, I. F., et al.
Pubblicazione: (2026)
di: Atasoy, I. F., et al.
Pubblicazione: (2026)
Moltbook Moderation: Uncovering Hidden Intent Through Multi-Turn Dialogue
di: Al-Lawati, Ali, et al.
Pubblicazione: (2026)
di: Al-Lawati, Ali, et al.
Pubblicazione: (2026)
GenAITEd Ghana: A First-of-Its-Kind Context-Aware and Curriculum-Aligned Conversational AI Agent for Teacher Education
di: Nyaaba, Matthew, et al.
Pubblicazione: (2025)
di: Nyaaba, Matthew, et al.
Pubblicazione: (2025)
Consensus Sampling for Safer Generative AI
di: Kalai, Adam Tauman, et al.
Pubblicazione: (2025)
di: Kalai, Adam Tauman, et al.
Pubblicazione: (2025)
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers
di: Achara, Akshit, et al.
Pubblicazione: (2025)
di: Achara, Akshit, et al.
Pubblicazione: (2025)
Breaking Free Transformer Models: Task-specific Context Attribution Promises Improved Generalizability Without Fine-tuning Pre-trained LLMs
di: Tytarenko, Stepan, et al.
Pubblicazione: (2024)
di: Tytarenko, Stepan, et al.
Pubblicazione: (2024)
Efficient LLM Moderation with Multi-Layer Latent Prototypes
di: Chrabąszcz, Maciej, et al.
Pubblicazione: (2025)
di: Chrabąszcz, Maciej, et al.
Pubblicazione: (2025)
PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media
di: Kachwala, Zoher, et al.
Pubblicazione: (2026)
di: Kachwala, Zoher, et al.
Pubblicazione: (2026)
Disentangling the Drivers of LLM Social Conformity: An Uncertainty-Moderated Dual-Process Mechanism
di: Zhong, Huixin, et al.
Pubblicazione: (2025)
di: Zhong, Huixin, et al.
Pubblicazione: (2025)
FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content Moderation
di: Ding, Zhihao, et al.
Pubblicazione: (2026)
di: Ding, Zhihao, et al.
Pubblicazione: (2026)
Unified Game Moderation: Soft-Prompting and LLM-Assisted Label Transfer for Resource-Efficient Toxicity Detection
di: Yang, Zachary, et al.
Pubblicazione: (2025)
di: Yang, Zachary, et al.
Pubblicazione: (2025)
Moderating Model Marketplaces: Platform Governance Puzzles for AI Intermediaries
di: Gorwa, Robert, et al.
Pubblicazione: (2023)
di: Gorwa, Robert, et al.
Pubblicazione: (2023)
Moderating Harm: Benchmarking Large Language Models for Cyberbullying Detection in YouTube Comments
di: Muminovic, Amel
Pubblicazione: (2025)
di: Muminovic, Amel
Pubblicazione: (2025)
AI Advocate: Educational Path to Transform Squads to the Future
di: Soares, Carla, et al.
Pubblicazione: (2026)
di: Soares, Carla, et al.
Pubblicazione: (2026)
Documenti analoghi
-
A Multi-Perspective Benchmark and Moderation Model for Evaluating Safety and Adversarial Robustness
di: Machlovi, Naseem, et al.
Pubblicazione: (2025) -
Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models
di: Ren, Juan, et al.
Pubblicazione: (2025) -
Redefining Elderly Care with Agentic AI: Challenges and Opportunities
di: Khalil, Ruhul Amin, et al.
Pubblicazione: (2025) -
VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare
di: Shetty, Anudeex, et al.
Pubblicazione: (2025) -
AI Benchmarks and Datasets for LLM Evaluation
di: Ivanov, Todor, et al.
Pubblicazione: (2024)