GuardReasoner: Towards Reasoning-based LLM Safeguards
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yue, Gao, Hongcheng, Zhai, Shengfang, He, Yufei, Xia, Jun, Hu, Zhengyu, Chen, Yulin, Yang, Xihong, Zhang, Jiaheng, Li, Stan Z., Xiong, Hui, Hooi, Bryan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio
by: Zhu, Zhenhao, et al.
Published: (2026)
by: Zhu, Zhenhao, et al.
Published: (2026)
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
by: Chen, Yulin, et al.
Published: (2026)
by: Chen, Yulin, et al.
Published: (2026)
GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
Securing LLM Agents Need Intent-to-Execution Integrity
by: Qu, Wenjie, et al.
Published: (2026)
by: Qu, Wenjie, et al.
Published: (2026)
Can Indirect Prompt Injection Attacks Be Detected and Removed?
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning
by: Guan, Shaowei, et al.
Published: (2025)
by: Guan, Shaowei, et al.
Published: (2025)
GLiGuard: Schema-Conditioned Classification for LLM Safeguard
by: Zaratiana, Urchade, et al.
Published: (2026)
by: Zaratiana, Urchade, et al.
Published: (2026)
TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense
by: Hao, Shuyang, et al.
Published: (2025)
by: Hao, Shuyang, et al.
Published: (2025)
IMMACULATE: A Practical LLM Auditing Framework via Verifiable Computation
by: Guo, Yanpei, et al.
Published: (2026)
by: Guo, Yanpei, et al.
Published: (2026)
Geneshift: Impact of different scenario shift on Jailbreaking LLM
by: Wu, Tianyi, et al.
Published: (2025)
by: Wu, Tianyi, et al.
Published: (2025)
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
by: Lin, Lixing, et al.
Published: (2026)
by: Lin, Lixing, et al.
Published: (2026)
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
ExtendAttack: Attacking Servers of LRMs via Extending Reasoning
by: Zhu, Zhenhao, et al.
Published: (2025)
by: Zhu, Zhenhao, et al.
Published: (2025)
Automated Phishing Detection Using URLs and Webpages
by: Wang, Huilin, et al.
Published: (2024)
by: Wang, Huilin, et al.
Published: (2024)
Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
by: Yang, Xianglin, et al.
Published: (2026)
by: Yang, Xianglin, et al.
Published: (2026)
FlipAttack: Jailbreak LLMs via Flipping
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing
by: Li, Yuexin, et al.
Published: (2026)
by: Li, Yuexin, et al.
Published: (2026)
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
by: Chen, Yulin, et al.
Published: (2024)
by: Chen, Yulin, et al.
Published: (2024)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
by: Chen, Yulin, et al.
Published: (2024)
by: Chen, Yulin, et al.
Published: (2024)
Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Efficient Input-level Backdoor Defense on Text-to-Image Synthesis via Neuron Activation Variation
by: Zhai, Shengfang, et al.
Published: (2025)
by: Zhai, Shengfang, et al.
Published: (2025)
LLM Safeguard is a Double-Edged Sword: Exploiting False Positives for Denial-of-Service Attacks
by: Zhang, Qingzhao, et al.
Published: (2024)
by: Zhang, Qingzhao, et al.
Published: (2024)
BadDLM: Backdooring Diffusion Language Models with Diverse Targets
by: Zhai, Shengfang, et al.
Published: (2026)
by: Zhai, Shengfang, et al.
Published: (2026)
PropGuard: Safeguarding LLM-MAS via Propagation-Aware Exploration and Remediation
by: Yan, Bingyu, et al.
Published: (2026)
by: Yan, Bingyu, et al.
Published: (2026)
ExpShield: Safeguarding Web Text from Unauthorized Crawling and LLM Exploitation
by: Liu, Ruixuan, et al.
Published: (2024)
by: Liu, Ruixuan, et al.
Published: (2024)
CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks
by: Li, Siyuan, et al.
Published: (2026)
by: Li, Siyuan, et al.
Published: (2026)
TraceGuard: Process-Guided Firewall against Reasoning Backdoors in Large Language Models
by: Guo, Zhen, et al.
Published: (2026)
by: Guo, Zhen, et al.
Published: (2026)
PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
by: Zhao, Jiawei, et al.
Published: (2025)
by: Zhao, Jiawei, et al.
Published: (2025)
DMark: Order-Agnostic Watermarking for Diffusion Large Language Models
by: Wu, Linyu, et al.
Published: (2025)
by: Wu, Linyu, et al.
Published: (2025)
MemPot: Defending Against Memory Extraction Attack with Optimized Honeypots
by: Wang, Yuhao, et al.
Published: (2026)
by: Wang, Yuhao, et al.
Published: (2026)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
by: Cao, Tri, et al.
Published: (2026)
by: Cao, Tri, et al.
Published: (2026)
CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
by: You, Wenhao, et al.
Published: (2025)
by: You, Wenhao, et al.
Published: (2025)
Life-Cycle Routing Vulnerabilities of LLM Router
by: Lin, Qiqi, et al.
Published: (2025)
by: Lin, Qiqi, et al.
Published: (2025)
SARSteer: Safeguarding Large Audio-Language Models via Safe-Ablated Refusal Steering
by: Lin, Weilin, et al.
Published: (2025)
by: Lin, Weilin, et al.
Published: (2025)
TransLinkGuard: Safeguarding Transformer Models Against Model Stealing in Edge Deployment
by: Li, Qinfeng, et al.
Published: (2024)
by: Li, Qinfeng, et al.
Published: (2024)
Autonomous Chain-of-Thought Distillation for Graph-Based Fraud Detection
by: Li, Yuan, et al.
Published: (2026)
by: Li, Yuan, et al.
Published: (2026)
Similar Items
-
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
by: Liu, Yue, et al.
Published: (2025) -
GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio
by: Zhu, Zhenhao, et al.
Published: (2026) -
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
by: Chen, Yulin, et al.
Published: (2026) -
GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners
by: Li, Haoran, et al.
Published: (2025) -
Securing LLM Agents Need Intent-to-Execution Integrity
by: Qu, Wenjie, et al.
Published: (2026)