PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
Fuente:
arXiv
Saved in:
| Main Authors: | Kumar, Priyanshu, Jain, Devansh, Yerukola, Akhila, Jiang, Liwei, Beniwal, Himanshu, Hartvigsen, Thomas, Sap, Maarten |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models
by: Jain, Devansh, et al.
Published: (2024)
by: Jain, Devansh, et al.
Published: (2024)
Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
by: Beniwal, Himanshu, et al.
Published: (2025)
by: Beniwal, Himanshu, et al.
Published: (2025)
Is the Pope Catholic? Yes, the Pope is Catholic. Generative Evaluation of Non-Literal Intent Resolution in LLMs
by: Yerukola, Akhila, et al.
Published: (2024)
by: Yerukola, Akhila, et al.
Published: (2024)
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models
by: Rao, Abhinav, et al.
Published: (2024)
by: Rao, Abhinav, et al.
Published: (2024)
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
by: Yerukola, Akhila, et al.
Published: (2025)
by: Yerukola, Akhila, et al.
Published: (2025)
Out of Style: RAG's Fragility to Linguistic Variation
by: Cao, Tianyu, et al.
Published: (2025)
by: Cao, Tianyu, et al.
Published: (2025)
Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities
by: Kim, Youngwoo, et al.
Published: (2025)
by: Kim, Youngwoo, et al.
Published: (2025)
Words Like Knives: Backstory-Personalized Modeling and Detection of Violent Communication
by: Shen, Jocelyn, et al.
Published: (2025)
by: Shen, Jocelyn, et al.
Published: (2025)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
by: Han, Seungju, et al.
Published: (2024)
by: Han, Seungju, et al.
Published: (2024)
Cross-lingual Editing in Multilingual Language Models
by: Beniwal, Himanshu, et al.
Published: (2024)
by: Beniwal, Himanshu, et al.
Published: (2024)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
by: Li, Jing-Jing, et al.
Published: (2024)
by: Li, Jing-Jing, et al.
Published: (2024)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
by: Beniwal, Himanshu, et al.
Published: (2026)
by: Beniwal, Himanshu, et al.
Published: (2026)
Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs
by: Beniwal, Himanshu, et al.
Published: (2025)
by: Beniwal, Himanshu, et al.
Published: (2025)
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
by: Zheng, Mingqian, et al.
Published: (2026)
by: Zheng, Mingqian, et al.
Published: (2026)
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
Social World Models
by: Zhou, Xuhui, et al.
Published: (2025)
by: Zhou, Xuhui, et al.
Published: (2025)
EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
by: Ghate, Kshitish, et al.
Published: (2025)
by: Ghate, Kshitish, et al.
Published: (2025)
COMMENTATOR: A Code-mixed Multilingual Text Annotation Framework
by: Sheth, Rajvee, et al.
Published: (2024)
by: Sheth, Rajvee, et al.
Published: (2024)
BingoGuard: LLM Content Moderation Tools with Risk Levels
by: Yin, Fan, et al.
Published: (2025)
by: Yin, Fan, et al.
Published: (2025)
Disparities in LLM Reasoning Accuracy and Explanations: A Case Study on African American English
by: Zhou, Runtao, et al.
Published: (2025)
by: Zhou, Runtao, et al.
Published: (2025)
DEPART: DEcomposing PARiTy across Multilingual LLMs
by: Uppadhyay, Manan, et al.
Published: (2026)
by: Uppadhyay, Manan, et al.
Published: (2026)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
by: Zhao, Yunhan, et al.
Published: (2026)
by: Zhao, Yunhan, et al.
Published: (2026)
CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications
by: Joshi, Raviraj, et al.
Published: (2025)
by: Joshi, Raviraj, et al.
Published: (2025)
Continually Self-Improving Language Models for Bariatric Surgery Question--Answering
by: Atri, Yash Kumar, et al.
Published: (2025)
by: Atri, Yash Kumar, et al.
Published: (2025)
MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety
by: Yang, Yahan, et al.
Published: (2025)
by: Yang, Yahan, et al.
Published: (2025)
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing
by: Sheth, Rajvee, et al.
Published: (2025)
by: Sheth, Rajvee, et al.
Published: (2025)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
by: Yadav, Ankit, et al.
Published: (2024)
by: Yadav, Ankit, et al.
Published: (2024)
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
by: Jiang, Liwei, et al.
Published: (2024)
by: Jiang, Liwei, et al.
Published: (2024)
Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation
by: Wróbel, Krzysztof, et al.
Published: (2026)
by: Wróbel, Krzysztof, et al.
Published: (2026)
LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
by: Tan, Leanne, et al.
Published: (2025)
by: Tan, Leanne, et al.
Published: (2025)
Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
by: Jiang, Liwei, et al.
Published: (2025)
by: Jiang, Liwei, et al.
Published: (2025)
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
by: Mendelsohn, Julia, et al.
Published: (2023)
by: Mendelsohn, Julia, et al.
Published: (2023)
Data Defenses Against Large Language Models
by: Agnew, William, et al.
Published: (2024)
by: Agnew, William, et al.
Published: (2024)
Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models across Modalities
by: Sheth, Rajvee, et al.
Published: (2025)
by: Sheth, Rajvee, et al.
Published: (2025)
TAXI: Evaluating Categorical Knowledge Editing for Language Models
by: Powell, Derek, et al.
Published: (2024)
by: Powell, Derek, et al.
Published: (2024)
MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention
by: Jha, Prince, et al.
Published: (2024)
by: Jha, Prince, et al.
Published: (2024)
UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages
by: Beniwal, Himanshu, et al.
Published: (2025)
by: Beniwal, Himanshu, et al.
Published: (2025)
Multilingual Information Retrieval with a Monolingual Knowledge Base
by: Zhuang, Yingying, et al.
Published: (2025)
by: Zhuang, Yingying, et al.
Published: (2025)
FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models
by: Fatehkia, Masoomali, et al.
Published: (2025)
by: Fatehkia, Masoomali, et al.
Published: (2025)
Model Editing with Graph-Based External Memory
by: Atri, Yash Kumar, et al.
Published: (2025)
by: Atri, Yash Kumar, et al.
Published: (2025)
Similar Items
-
PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models
by: Jain, Devansh, et al.
Published: (2024) -
Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
by: Beniwal, Himanshu, et al.
Published: (2025) -
Is the Pope Catholic? Yes, the Pope is Catholic. Generative Evaluation of Non-Literal Intent Resolution in LLMs
by: Yerukola, Akhila, et al.
Published: (2024) -
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models
by: Rao, Abhinav, et al.
Published: (2024) -
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
by: Yerukola, Akhila, et al.
Published: (2025)