MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Yahan, Dan, Soham, Li, Shuo, Roth, Dan, Lee, Insup |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Benchmarking LLM Guardrails in Handling Multilingual Toxicity
by: Yang, Yahan, et al.
Published: (2024)
by: Yang, Yahan, et al.
Published: (2024)
On the Calibration of Multilingual Question Answering LLMs
by: Yang, Yahan, et al.
Published: (2023)
by: Yang, Yahan, et al.
Published: (2023)
DuoGuard: A Two-Player RL-Driven Framework for Multilingual LLM Guardrails
by: Deng, Yihe, et al.
Published: (2025)
by: Deng, Yihe, et al.
Published: (2025)
ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails
by: Wang, Yan, et al.
Published: (2026)
by: Wang, Yan, et al.
Published: (2026)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
by: Zhao, Yunhan, et al.
Published: (2026)
by: Zhao, Yunhan, et al.
Published: (2026)
IOLBENCH: Benchmarking LLMs on Linguistic Reasoning
by: Goyal, Satyam, et al.
Published: (2025)
by: Goyal, Satyam, et al.
Published: (2025)
Bag of Tricks for Subverting Reasoning-based Safety Guardrails
by: Chen, Shuo, et al.
Published: (2025)
by: Chen, Shuo, et al.
Published: (2025)
Leveraging LLM For Synchronizing Information Across Multilingual Tables
by: Khincha, Siddharth, et al.
Published: (2025)
by: Khincha, Siddharth, et al.
Published: (2025)
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
by: Oh, Sejoon, et al.
Published: (2024)
by: Oh, Sejoon, et al.
Published: (2024)
Reasoning is about giving reasons
by: Shah, Krunal, et al.
Published: (2025)
by: Shah, Krunal, et al.
Published: (2025)
Multilingual Needle in a Haystack: Investigating Long-Context Behavior of Multilingual Large Language Models
by: Hengle, Amey, et al.
Published: (2024)
by: Hengle, Amey, et al.
Published: (2024)
Guard Vector: Beyond English LLM Guardrails with Task-Vector Composition and Streaming-Aware Prefix SFT
by: Lee, Wonhyuk, et al.
Published: (2025)
by: Lee, Wonhyuk, et al.
Published: (2025)
RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
by: She, Yining, et al.
Published: (2025)
by: She, Yining, et al.
Published: (2025)
ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
by: Aswal, Darpan, et al.
Published: (2025)
by: Aswal, Darpan, et al.
Published: (2025)
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
by: Zhu, Boyu, et al.
Published: (2025)
by: Zhu, Boyu, et al.
Published: (2025)
PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
by: Kumar, Priyanshu, et al.
Published: (2025)
by: Kumar, Priyanshu, et al.
Published: (2025)
LLM-Symbolic Integration for Robust Temporal Tabular Reasoning
by: Kulkarni, Atharv, et al.
Published: (2025)
by: Kulkarni, Atharv, et al.
Published: (2025)
SentGuard: Sentence-Level Streaming Guardrails for Large Language Models
by: Yu, Jiaqi, et al.
Published: (2026)
by: Yu, Jiaqi, et al.
Published: (2026)
Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2025)
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2025)
CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications
by: Joshi, Raviraj, et al.
Published: (2025)
by: Joshi, Raviraj, et al.
Published: (2025)
TRAQ: Trustworthy Retrieval Augmented Question Answering via Conformal Prediction
by: Li, Shuo, et al.
Published: (2023)
by: Li, Shuo, et al.
Published: (2023)
Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
by: Giarrusso, Francesco, et al.
Published: (2025)
by: Giarrusso, Francesco, et al.
Published: (2025)
Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
XTRUST: On the Multilingual Trustworthiness of Large Language Models
by: Li, Yahan, et al.
Published: (2024)
by: Li, Yahan, et al.
Published: (2024)
CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
by: An, Heajun, et al.
Published: (2026)
by: An, Heajun, et al.
Published: (2026)
MultiZebraLogic: A Multilingual Logical Reasoning Benchmark
by: Bruun, Sofie Helene, et al.
Published: (2025)
by: Bruun, Sofie Helene, et al.
Published: (2025)
No Universal Prompt: Unifying Reasoning through Adaptive Prompting for Temporal Table Reasoning
by: Rajgaria, Abhishek, et al.
Published: (2025)
by: Rajgaria, Abhishek, et al.
Published: (2025)
WebGuard: Building a Generalizable Guardrail for Web Agents
by: Zheng, Boyuan, et al.
Published: (2025)
by: Zheng, Boyuan, et al.
Published: (2025)
TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts
by: Chu, Hua-Rong, et al.
Published: (2026)
by: Chu, Hua-Rong, et al.
Published: (2026)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding
by: Yoo, Haneul, et al.
Published: (2024)
by: Yoo, Haneul, et al.
Published: (2024)
SocREval: Large Language Models with the Socratic Method for Reference-Free Reasoning Evaluation
by: He, Hangfeng, et al.
Published: (2023)
by: He, Hangfeng, et al.
Published: (2023)
NAVCON: A Cognitively Inspired and Linguistically Grounded Corpus for Vision and Language Navigation
by: Wanchoo, Karan, et al.
Published: (2024)
by: Wanchoo, Karan, et al.
Published: (2024)
Thinking Machines: A Survey of LLM based Reasoning Strategies
by: Bandyopadhyay, Dibyanayan, et al.
Published: (2025)
by: Bandyopadhyay, Dibyanayan, et al.
Published: (2025)
H-STAR: LLM-driven Hybrid SQL-Text Adaptive Reasoning on Tables
by: Abhyankar, Nikhil, et al.
Published: (2024)
by: Abhyankar, Nikhil, et al.
Published: (2024)
ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments
by: Wang, Yuquan, et al.
Published: (2025)
by: Wang, Yuquan, et al.
Published: (2025)
Conflicts in Texts: Data, Implications and Challenges
by: Liu, Siyi, et al.
Published: (2025)
by: Liu, Siyi, et al.
Published: (2025)
A Lightweight Explainable Guardrail for Prompt Safety
by: Islam, Md Asiful, et al.
Published: (2026)
by: Islam, Md Asiful, et al.
Published: (2026)
Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails
by: Ghosh, Shaona, et al.
Published: (2025)
by: Ghosh, Shaona, et al.
Published: (2025)
CREST: Universal Safety Guardrails Through Cluster-Guided Cross-Lingual Transfer
by: Bansal, Lavish, et al.
Published: (2025)
by: Bansal, Lavish, et al.
Published: (2025)
Similar Items
-
Benchmarking LLM Guardrails in Handling Multilingual Toxicity
by: Yang, Yahan, et al.
Published: (2024) -
On the Calibration of Multilingual Question Answering LLMs
by: Yang, Yahan, et al.
Published: (2023) -
DuoGuard: A Two-Player RL-Driven Framework for Multilingual LLM Guardrails
by: Deng, Yihe, et al.
Published: (2025) -
ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails
by: Wang, Yan, et al.
Published: (2026) -
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
by: Zhao, Yunhan, et al.
Published: (2026)