PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
Fuente:
arXiv
Guardado en:
| Autores principales: | Kumar, Priyanshu, Jain, Devansh, Yerukola, Akhila, Jiang, Liwei, Beniwal, Himanshu, Hartvigsen, Thomas, Sap, Maarten |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models
por: Jain, Devansh, et al.
Publicado: (2024)
por: Jain, Devansh, et al.
Publicado: (2024)
Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
por: Beniwal, Himanshu, et al.
Publicado: (2025)
por: Beniwal, Himanshu, et al.
Publicado: (2025)
Is the Pope Catholic? Yes, the Pope is Catholic. Generative Evaluation of Non-Literal Intent Resolution in LLMs
por: Yerukola, Akhila, et al.
Publicado: (2024)
por: Yerukola, Akhila, et al.
Publicado: (2024)
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models
por: Rao, Abhinav, et al.
Publicado: (2024)
por: Rao, Abhinav, et al.
Publicado: (2024)
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
por: Yerukola, Akhila, et al.
Publicado: (2025)
por: Yerukola, Akhila, et al.
Publicado: (2025)
Out of Style: RAG's Fragility to Linguistic Variation
por: Cao, Tianyu, et al.
Publicado: (2025)
por: Cao, Tianyu, et al.
Publicado: (2025)
Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities
por: Kim, Youngwoo, et al.
Publicado: (2025)
por: Kim, Youngwoo, et al.
Publicado: (2025)
Words Like Knives: Backstory-Personalized Modeling and Detection of Violent Communication
por: Shen, Jocelyn, et al.
Publicado: (2025)
por: Shen, Jocelyn, et al.
Publicado: (2025)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
por: Han, Seungju, et al.
Publicado: (2024)
por: Han, Seungju, et al.
Publicado: (2024)
Cross-lingual Editing in Multilingual Language Models
por: Beniwal, Himanshu, et al.
Publicado: (2024)
por: Beniwal, Himanshu, et al.
Publicado: (2024)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
por: Li, Jing-Jing, et al.
Publicado: (2024)
por: Li, Jing-Jing, et al.
Publicado: (2024)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
por: Beniwal, Himanshu, et al.
Publicado: (2026)
por: Beniwal, Himanshu, et al.
Publicado: (2026)
Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs
por: Beniwal, Himanshu, et al.
Publicado: (2025)
por: Beniwal, Himanshu, et al.
Publicado: (2025)
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
por: Zheng, Mingqian, et al.
Publicado: (2026)
por: Zheng, Mingqian, et al.
Publicado: (2026)
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
Social World Models
por: Zhou, Xuhui, et al.
Publicado: (2025)
por: Zhou, Xuhui, et al.
Publicado: (2025)
EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
por: Ghate, Kshitish, et al.
Publicado: (2025)
por: Ghate, Kshitish, et al.
Publicado: (2025)
COMMENTATOR: A Code-mixed Multilingual Text Annotation Framework
por: Sheth, Rajvee, et al.
Publicado: (2024)
por: Sheth, Rajvee, et al.
Publicado: (2024)
BingoGuard: LLM Content Moderation Tools with Risk Levels
por: Yin, Fan, et al.
Publicado: (2025)
por: Yin, Fan, et al.
Publicado: (2025)
Disparities in LLM Reasoning Accuracy and Explanations: A Case Study on African American English
por: Zhou, Runtao, et al.
Publicado: (2025)
por: Zhou, Runtao, et al.
Publicado: (2025)
DEPART: DEcomposing PARiTy across Multilingual LLMs
por: Uppadhyay, Manan, et al.
Publicado: (2026)
por: Uppadhyay, Manan, et al.
Publicado: (2026)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
por: Zhao, Yunhan, et al.
Publicado: (2026)
por: Zhao, Yunhan, et al.
Publicado: (2026)
CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications
por: Joshi, Raviraj, et al.
Publicado: (2025)
por: Joshi, Raviraj, et al.
Publicado: (2025)
Continually Self-Improving Language Models for Bariatric Surgery Question--Answering
por: Atri, Yash Kumar, et al.
Publicado: (2025)
por: Atri, Yash Kumar, et al.
Publicado: (2025)
MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety
por: Yang, Yahan, et al.
Publicado: (2025)
por: Yang, Yahan, et al.
Publicado: (2025)
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing
por: Sheth, Rajvee, et al.
Publicado: (2025)
por: Sheth, Rajvee, et al.
Publicado: (2025)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
por: Yadav, Ankit, et al.
Publicado: (2024)
por: Yadav, Ankit, et al.
Publicado: (2024)
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
por: Jiang, Liwei, et al.
Publicado: (2024)
por: Jiang, Liwei, et al.
Publicado: (2024)
Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation
por: Wróbel, Krzysztof, et al.
Publicado: (2026)
por: Wróbel, Krzysztof, et al.
Publicado: (2026)
LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
por: Tan, Leanne, et al.
Publicado: (2025)
por: Tan, Leanne, et al.
Publicado: (2025)
Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
por: Jiang, Liwei, et al.
Publicado: (2025)
por: Jiang, Liwei, et al.
Publicado: (2025)
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
por: Mendelsohn, Julia, et al.
Publicado: (2023)
por: Mendelsohn, Julia, et al.
Publicado: (2023)
Data Defenses Against Large Language Models
por: Agnew, William, et al.
Publicado: (2024)
por: Agnew, William, et al.
Publicado: (2024)
Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models across Modalities
por: Sheth, Rajvee, et al.
Publicado: (2025)
por: Sheth, Rajvee, et al.
Publicado: (2025)
TAXI: Evaluating Categorical Knowledge Editing for Language Models
por: Powell, Derek, et al.
Publicado: (2024)
por: Powell, Derek, et al.
Publicado: (2024)
MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention
por: Jha, Prince, et al.
Publicado: (2024)
por: Jha, Prince, et al.
Publicado: (2024)
UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages
por: Beniwal, Himanshu, et al.
Publicado: (2025)
por: Beniwal, Himanshu, et al.
Publicado: (2025)
Multilingual Information Retrieval with a Monolingual Knowledge Base
por: Zhuang, Yingying, et al.
Publicado: (2025)
por: Zhuang, Yingying, et al.
Publicado: (2025)
FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models
por: Fatehkia, Masoomali, et al.
Publicado: (2025)
por: Fatehkia, Masoomali, et al.
Publicado: (2025)
Model Editing with Graph-Based External Memory
por: Atri, Yash Kumar, et al.
Publicado: (2025)
por: Atri, Yash Kumar, et al.
Publicado: (2025)
Ejemplares similares
-
PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models
por: Jain, Devansh, et al.
Publicado: (2024) -
Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
por: Beniwal, Himanshu, et al.
Publicado: (2025) -
Is the Pope Catholic? Yes, the Pope is Catholic. Generative Evaluation of Non-Literal Intent Resolution in LLMs
por: Yerukola, Akhila, et al.
Publicado: (2024) -
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models
por: Rao, Abhinav, et al.
Publicado: (2024) -
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
por: Yerukola, Akhila, et al.
Publicado: (2025)