Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
Fuente:
arXiv
Saved in:
| Main Authors: | Kumar, Shanu, Kholkar, Gauri, Mendke, Saish, Sadana, Anubhav, Agrawal, Parag, Dandapat, Sandipan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Zero-shot Chain of Thought Prompting via Uncertainty-Guided Strategy Selection
by: Kumar, Shanu, et al.
Published: (2024)
by: Kumar, Shanu, et al.
Published: (2024)
CAPTURE: Context-Aware Prompt Injection Testing and Robustness Enhancement
by: Kholkar, Gauri, et al.
Published: (2025)
by: Kholkar, Gauri, et al.
Published: (2025)
Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
by: Mittal, Avni, et al.
Published: (2026)
by: Mittal, Avni, et al.
Published: (2026)
Policy-as-Prompt: Turning AI Governance Rules into Guardrails for AI Agents
by: Kholkar, Gauri, et al.
Published: (2025)
by: Kholkar, Gauri, et al.
Published: (2025)
SAGE: A Generic Framework for LLM Safety Evaluation
by: Jindal, Madhur, et al.
Published: (2025)
by: Jindal, Madhur, et al.
Published: (2025)
LLM Safety for Children
by: Rath, Prasanjit, et al.
Published: (2025)
by: Rath, Prasanjit, et al.
Published: (2025)
Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations
by: Banerjee, Somnath, et al.
Published: (2025)
by: Banerjee, Somnath, et al.
Published: (2025)
Navigating the Cultural Kaleidoscope: A Hitchhiker's Guide to Sensitivity in Large Language Models
by: Banerjee, Somnath, et al.
Published: (2024)
by: Banerjee, Somnath, et al.
Published: (2024)
Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment
by: Vemula, Saketh Reddy, et al.
Published: (2025)
by: Vemula, Saketh Reddy, et al.
Published: (2025)
READ: Reinforcement-based Adversarial Learning for Text Classification with Limited Labeled Data
by: Sharma, Rohit, et al.
Published: (2025)
by: Sharma, Rohit, et al.
Published: (2025)
PADBen: A Comprehensive Benchmark for Evaluating AI Text Detectors Against Paraphrase Attacks
by: Zha, Yiwei, et al.
Published: (2025)
by: Zha, Yiwei, et al.
Published: (2025)
Context-Aware Content Moderation for German Newspaper Comments
by: Krejca, Felix, et al.
Published: (2025)
by: Krejca, Felix, et al.
Published: (2025)
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
by: Zhang, Hongbin, et al.
Published: (2025)
by: Zhang, Hongbin, et al.
Published: (2025)
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
by: Mendu, Sai Krishna, et al.
Published: (2025)
by: Mendu, Sai Krishna, et al.
Published: (2025)
HiFACTMix: A Code-Mixed Benchmark and Graph-Aware Model for EvidenceBased Political Claim Verification in Hinglish
by: Thakur, Rakesh, et al.
Published: (2025)
by: Thakur, Rakesh, et al.
Published: (2025)
Beyond Next Word Prediction: Developing Comprehensive Evaluation Frameworks for measuring LLM performance on real world applications
by: Agrawal, Vishakha, et al.
Published: (2025)
by: Agrawal, Vishakha, et al.
Published: (2025)
LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation
by: Park, Junyeong, et al.
Published: (2025)
by: Park, Junyeong, et al.
Published: (2025)
Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
MUGC: Machine Generated versus User Generated Content Detection
by: Xie, Yaqi, et al.
Published: (2024)
by: Xie, Yaqi, et al.
Published: (2024)
Evaluating the Diversity and Quality of LLM Generated Content
by: Shypula, Alexander, et al.
Published: (2025)
by: Shypula, Alexander, et al.
Published: (2025)
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
by: Kreutzer, Julia, et al.
Published: (2025)
by: Kreutzer, Julia, et al.
Published: (2025)
IPS: In-Prompt Process Supervision for Short Video Content Moderation
by: Liu, Mingchao, et al.
Published: (2024)
by: Liu, Mingchao, et al.
Published: (2024)
Meta-Cultural Competence: Climbing the Right Hill of Cultural Awareness
by: Saha, Sougata, et al.
Published: (2025)
by: Saha, Sougata, et al.
Published: (2025)
Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning
by: Kim, Minwu, et al.
Published: (2025)
by: Kim, Minwu, et al.
Published: (2025)
Class-RAG: Real-Time Content Moderation with Retrieval Augmented Generation
by: Chen, Jianfa, et al.
Published: (2024)
by: Chen, Jianfa, et al.
Published: (2024)
Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation
by: Wróbel, Krzysztof, et al.
Published: (2026)
by: Wróbel, Krzysztof, et al.
Published: (2026)
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation
by: Schwartz, Daniel, et al.
Published: (2025)
by: Schwartz, Daniel, et al.
Published: (2025)
Boosting Zero-Shot Crosslingual Performance using LLM-Based Augmentations with Effective Data Selection
by: Fazili, Barah, et al.
Published: (2024)
by: Fazili, Barah, et al.
Published: (2024)
No-Worse Context-Aware Decoding: Preventing Neutral Regression in Context-Conditioned Generation
by: Tao, Yufei, et al.
Published: (2026)
by: Tao, Yufei, et al.
Published: (2026)
Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation
by: Zhang, Shutong, et al.
Published: (2026)
by: Zhang, Shutong, et al.
Published: (2026)
Evaluating Cultural Awareness of LLMs for Yoruba, Malayalam, and English
by: Dawson, Fiifi, et al.
Published: (2024)
by: Dawson, Fiifi, et al.
Published: (2024)
Enhancing Cache-Augmented Generation (CAG) with Adaptive Contextual Compression for Scalable Knowledge Integration
by: Agrawal, Rishabh, et al.
Published: (2025)
by: Agrawal, Rishabh, et al.
Published: (2025)
Probing Association Biases in LLM Moderation Over-Sensitivity
by: Wang, Yuxin, et al.
Published: (2025)
by: Wang, Yuxin, et al.
Published: (2025)
Structure-BiEval: A Self-Supervised, Dual-Track Framework for Decoupling Structure and Content in LLM Evaluation for Web Information Systems
by: Zhao, Boxiang, et al.
Published: (2026)
by: Zhao, Boxiang, et al.
Published: (2026)
CultureBank: An Online Community-Driven Knowledge Base Towards Culturally Aware Language Technologies
by: Shi, Weiyan, et al.
Published: (2024)
by: Shi, Weiyan, et al.
Published: (2024)
LionGuard: Building a Contextualized Moderation Classifier to Tackle Localized Unsafe Content
by: Foo, Jessica, et al.
Published: (2024)
by: Foo, Jessica, et al.
Published: (2024)
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation
by: Zhuang, Jun, et al.
Published: (2025)
by: Zhuang, Jun, et al.
Published: (2025)
Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators
by: Cao, Yang Trista, et al.
Published: (2023)
by: Cao, Yang Trista, et al.
Published: (2023)
Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs
by: Han, HyoJung, et al.
Published: (2025)
by: Han, HyoJung, et al.
Published: (2025)
MTQ-Eval: Multilingual Text Quality Evaluation for Language Models
by: Pokharel, Rhitabrat, et al.
Published: (2025)
by: Pokharel, Rhitabrat, et al.
Published: (2025)
Similar Items
-
Enhancing Zero-shot Chain of Thought Prompting via Uncertainty-Guided Strategy Selection
by: Kumar, Shanu, et al.
Published: (2024) -
CAPTURE: Context-Aware Prompt Injection Testing and Robustness Enhancement
by: Kholkar, Gauri, et al.
Published: (2025) -
Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
by: Mittal, Avni, et al.
Published: (2026) -
Policy-as-Prompt: Turning AI Governance Rules into Guardrails for AI Agents
by: Kholkar, Gauri, et al.
Published: (2025) -
SAGE: A Generic Framework for LLM Safety Evaluation
by: Jindal, Madhur, et al.
Published: (2025)