PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kumar, Priyanshu, Jain, Devansh, Yerukola, Akhila, Jiang, Liwei, Beniwal, Himanshu, Hartvigsen, Thomas, Sap, Maarten |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models
von: Jain, Devansh, et al.
Veröffentlicht: (2024)
von: Jain, Devansh, et al.
Veröffentlicht: (2024)
Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
Is the Pope Catholic? Yes, the Pope is Catholic. Generative Evaluation of Non-Literal Intent Resolution in LLMs
von: Yerukola, Akhila, et al.
Veröffentlicht: (2024)
von: Yerukola, Akhila, et al.
Veröffentlicht: (2024)
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models
von: Rao, Abhinav, et al.
Veröffentlicht: (2024)
von: Rao, Abhinav, et al.
Veröffentlicht: (2024)
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
von: Yerukola, Akhila, et al.
Veröffentlicht: (2025)
von: Yerukola, Akhila, et al.
Veröffentlicht: (2025)
Out of Style: RAG's Fragility to Linguistic Variation
von: Cao, Tianyu, et al.
Veröffentlicht: (2025)
von: Cao, Tianyu, et al.
Veröffentlicht: (2025)
Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities
von: Kim, Youngwoo, et al.
Veröffentlicht: (2025)
von: Kim, Youngwoo, et al.
Veröffentlicht: (2025)
Words Like Knives: Backstory-Personalized Modeling and Detection of Violent Communication
von: Shen, Jocelyn, et al.
Veröffentlicht: (2025)
von: Shen, Jocelyn, et al.
Veröffentlicht: (2025)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
von: Han, Seungju, et al.
Veröffentlicht: (2024)
von: Han, Seungju, et al.
Veröffentlicht: (2024)
Cross-lingual Editing in Multilingual Language Models
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2024)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2024)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
von: Li, Jing-Jing, et al.
Veröffentlicht: (2024)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2024)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2026)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2026)
Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
von: Zheng, Mingqian, et al.
Veröffentlicht: (2026)
von: Zheng, Mingqian, et al.
Veröffentlicht: (2026)
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
Social World Models
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
von: Ghate, Kshitish, et al.
Veröffentlicht: (2025)
von: Ghate, Kshitish, et al.
Veröffentlicht: (2025)
COMMENTATOR: A Code-mixed Multilingual Text Annotation Framework
von: Sheth, Rajvee, et al.
Veröffentlicht: (2024)
von: Sheth, Rajvee, et al.
Veröffentlicht: (2024)
BingoGuard: LLM Content Moderation Tools with Risk Levels
von: Yin, Fan, et al.
Veröffentlicht: (2025)
von: Yin, Fan, et al.
Veröffentlicht: (2025)
Disparities in LLM Reasoning Accuracy and Explanations: A Case Study on African American English
von: Zhou, Runtao, et al.
Veröffentlicht: (2025)
von: Zhou, Runtao, et al.
Veröffentlicht: (2025)
DEPART: DEcomposing PARiTy across Multilingual LLMs
von: Uppadhyay, Manan, et al.
Veröffentlicht: (2026)
von: Uppadhyay, Manan, et al.
Veröffentlicht: (2026)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
von: Zhao, Yunhan, et al.
Veröffentlicht: (2026)
von: Zhao, Yunhan, et al.
Veröffentlicht: (2026)
CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications
von: Joshi, Raviraj, et al.
Veröffentlicht: (2025)
von: Joshi, Raviraj, et al.
Veröffentlicht: (2025)
Continually Self-Improving Language Models for Bariatric Surgery Question--Answering
von: Atri, Yash Kumar, et al.
Veröffentlicht: (2025)
von: Atri, Yash Kumar, et al.
Veröffentlicht: (2025)
MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety
von: Yang, Yahan, et al.
Veröffentlicht: (2025)
von: Yang, Yahan, et al.
Veröffentlicht: (2025)
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing
von: Sheth, Rajvee, et al.
Veröffentlicht: (2025)
von: Sheth, Rajvee, et al.
Veröffentlicht: (2025)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
von: Yadav, Ankit, et al.
Veröffentlicht: (2024)
von: Yadav, Ankit, et al.
Veröffentlicht: (2024)
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
von: Jiang, Liwei, et al.
Veröffentlicht: (2024)
von: Jiang, Liwei, et al.
Veröffentlicht: (2024)
Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation
von: Wróbel, Krzysztof, et al.
Veröffentlicht: (2026)
von: Wróbel, Krzysztof, et al.
Veröffentlicht: (2026)
LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
von: Tan, Leanne, et al.
Veröffentlicht: (2025)
von: Tan, Leanne, et al.
Veröffentlicht: (2025)
Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
von: Jiang, Liwei, et al.
Veröffentlicht: (2025)
von: Jiang, Liwei, et al.
Veröffentlicht: (2025)
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
von: Mendelsohn, Julia, et al.
Veröffentlicht: (2023)
von: Mendelsohn, Julia, et al.
Veröffentlicht: (2023)
Data Defenses Against Large Language Models
von: Agnew, William, et al.
Veröffentlicht: (2024)
von: Agnew, William, et al.
Veröffentlicht: (2024)
Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models across Modalities
von: Sheth, Rajvee, et al.
Veröffentlicht: (2025)
von: Sheth, Rajvee, et al.
Veröffentlicht: (2025)
TAXI: Evaluating Categorical Knowledge Editing for Language Models
von: Powell, Derek, et al.
Veröffentlicht: (2024)
von: Powell, Derek, et al.
Veröffentlicht: (2024)
MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention
von: Jha, Prince, et al.
Veröffentlicht: (2024)
von: Jha, Prince, et al.
Veröffentlicht: (2024)
UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
Multilingual Information Retrieval with a Monolingual Knowledge Base
von: Zhuang, Yingying, et al.
Veröffentlicht: (2025)
von: Zhuang, Yingying, et al.
Veröffentlicht: (2025)
FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models
von: Fatehkia, Masoomali, et al.
Veröffentlicht: (2025)
von: Fatehkia, Masoomali, et al.
Veröffentlicht: (2025)
Model Editing with Graph-Based External Memory
von: Atri, Yash Kumar, et al.
Veröffentlicht: (2025)
von: Atri, Yash Kumar, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models
von: Jain, Devansh, et al.
Veröffentlicht: (2024) -
Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025) -
Is the Pope Catholic? Yes, the Pope is Catholic. Generative Evaluation of Non-Literal Intent Resolution in LLMs
von: Yerukola, Akhila, et al.
Veröffentlicht: (2024) -
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models
von: Rao, Abhinav, et al.
Veröffentlicht: (2024) -
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
von: Yerukola, Akhila, et al.
Veröffentlicht: (2025)