Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Mintong, Chen, Zhaorun, Xu, Chejian, Zhang, Jiawei, Guo, Chengquan, Pan, Minzhou, Revilla, Ivan, Sun, Yu, Li, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning
by: Chen, Zhaorun, et al.
Published: (2025)
by: Chen, Zhaorun, et al.
Published: (2025)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
by: Zhao, Yunhan, et al.
Published: (2026)
by: Zhao, Yunhan, et al.
Published: (2026)
SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
by: Xu, Peiyang, et al.
Published: (2025)
by: Xu, Peiyang, et al.
Published: (2025)
AdvWave: Stealthy Adversarial Jailbreak Attack against Large Audio-Language Models
by: Kang, Mintong, et al.
Published: (2024)
by: Kang, Mintong, et al.
Published: (2024)
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
by: Xu, Chejian, et al.
Published: (2024)
by: Xu, Chejian, et al.
Published: (2024)
PoseGuard: Pose-Guided Generation with Safety Guardrails
by: Wang, Kongxin, et al.
Published: (2025)
by: Wang, Kongxin, et al.
Published: (2025)
MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models
by: Xu, Chejian, et al.
Published: (2025)
by: Xu, Chejian, et al.
Published: (2025)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
by: Liao, Zeyi, et al.
Published: (2024)
by: Liao, Zeyi, et al.
Published: (2024)
GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio
by: Zhu, Zhenhao, et al.
Published: (2026)
by: Zhu, Zhenhao, et al.
Published: (2026)
DiffAttack: Evasion Attacks Against Diffusion-Based Adversarial Purification
by: Kang, Mintong, et al.
Published: (2023)
by: Kang, Mintong, et al.
Published: (2023)
Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations
by: Chen, Zhaorun, et al.
Published: (2024)
by: Chen, Zhaorun, et al.
Published: (2024)
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
by: Li, Xu, et al.
Published: (2026)
by: Li, Xu, et al.
Published: (2026)
LPG: Balancing Efficiency and Policy Reasoning in Latent Policy Guardrails
by: Li, Nanxi, et al.
Published: (2026)
by: Li, Nanxi, et al.
Published: (2026)
Bag of Tricks for Subverting Reasoning-based Safety Guardrails
by: Chen, Shuo, et al.
Published: (2025)
by: Chen, Shuo, et al.
Published: (2025)
Hoist with His Own Petard: Inducing Guardrails to Facilitate Denial-of-Service Attacks on Retrieval-Augmented Generation of LLMs
by: Suo, Pan, et al.
Published: (2025)
by: Suo, Pan, et al.
Published: (2025)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysis
by: Kang, Mintong, et al.
Published: (2025)
by: Kang, Mintong, et al.
Published: (2025)
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
by: Wen, Xiaofei, et al.
Published: (2025)
by: Wen, Xiaofei, et al.
Published: (2025)
Guardrails as Infrastructure: Policy-First Control for Tool-Orchestrated Workflows
by: Sigdel, Akshey, et al.
Published: (2026)
by: Sigdel, Akshey, et al.
Published: (2026)
GLiNER Guard: Unified Encoder Family for Production LLM Safety and Privacy
by: Minko, Bogdan, et al.
Published: (2026)
by: Minko, Bogdan, et al.
Published: (2026)
QGuard:Question-based Zero-shot Guard for Multi-modal LLM Safety
by: Lee, Taegyeong, et al.
Published: (2025)
by: Lee, Taegyeong, et al.
Published: (2025)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025)
by: Yuan, Lingzhi, et al.
Published: (2025)
Novel Blockchain-based Protocols for Electronic Voting and Auctions
by: Lin, Zhaorun
Published: (2025)
by: Lin, Zhaorun
Published: (2025)
PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
by: Zhao, Jiawei, et al.
Published: (2025)
by: Zhao, Jiawei, et al.
Published: (2025)
Guard-GBDT: Efficient Privacy-Preserving Approximated GBDT Training on Vertical Dataset
by: Song, Anxiao, et al.
Published: (2025)
by: Song, Anxiao, et al.
Published: (2025)
TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts
by: Chu, Hua-Rong, et al.
Published: (2026)
by: Chu, Hua-Rong, et al.
Published: (2026)
COMMIT: Certifying Robustness of Multi-Sensor Fusion Systems against Semantic Attacks
by: Huang, Zijian, et al.
Published: (2024)
by: Huang, Zijian, et al.
Published: (2024)
No Free Lunch with Guardrails
by: Kumar, Divyanshu, et al.
Published: (2025)
by: Kumar, Divyanshu, et al.
Published: (2025)
SGuard-v1: Safety Guardrail for Large Language Models
by: Lee, JoonHo, et al.
Published: (2025)
by: Lee, JoonHo, et al.
Published: (2025)
Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
by: Young, Richard J.
Published: (2025)
by: Young, Richard J.
Published: (2025)
ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models
by: Jin, Weifei, et al.
Published: (2025)
by: Jin, Weifei, et al.
Published: (2025)
Peering Behind the Shield: Guardrail Identification in Large Language Models
by: Yang, Ziqing, et al.
Published: (2025)
by: Yang, Ziqing, et al.
Published: (2025)
The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search
by: Wei, Rongzhe, et al.
Published: (2025)
by: Wei, Rongzhe, et al.
Published: (2025)
OpenGuardrails: A Configurable, Unified, and Scalable Guardrails Platform for Large Language Models
by: Wang, Thomas, et al.
Published: (2025)
by: Wang, Thomas, et al.
Published: (2025)
Triaging Threats to Specialized Guardrails
by: Mo, Wenjie Jacky, et al.
Published: (2026)
by: Mo, Wenjie Jacky, et al.
Published: (2026)
Provably Secure Agent Guardrail
by: Wu, Benlong, et al.
Published: (2026)
by: Wu, Benlong, et al.
Published: (2026)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
by: Shen, Guobin, et al.
Published: (2025)
by: Shen, Guobin, et al.
Published: (2025)
AgentGuard: Repurposing Agentic Orchestrator for Safety Evaluation of Tool Orchestration
by: Chen, Jizhou, et al.
Published: (2025)
by: Chen, Jizhou, et al.
Published: (2025)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
by: Hsiung, Lei, et al.
Published: (2025)
by: Hsiung, Lei, et al.
Published: (2025)
Similar Items
-
ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning
by: Chen, Zhaorun, et al.
Published: (2025) -
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
by: Zhao, Yunhan, et al.
Published: (2026) -
SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
by: Xu, Peiyang, et al.
Published: (2025) -
AdvWave: Stealthy Adversarial Jailbreak Attack against Large Audio-Language Models
by: Kang, Mintong, et al.
Published: (2024) -
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
by: Xu, Chejian, et al.
Published: (2024)