GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Xiang, Zhen, Zheng, Linzhi, Li, Yanjie, Hong, Junyuan, Li, Qinbin, Xie, Han, Zhang, Jiawei, Xiong, Zidi, Xie, Chulin, Yang, Carl, Song, Dawn, Li, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GuardReasoner: Towards Reasoning-based LLM Safeguards
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
by: Chen, Zhaorun, et al.
Published: (2024)
by: Chen, Zhaorun, et al.
Published: (2024)
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
by: Chen, Yulin, et al.
Published: (2026)
by: Chen, Yulin, et al.
Published: (2026)
TrinityGuard: A Unified Framework for Safeguarding Multi-Agent Systems
by: Wang, Kai, et al.
Published: (2026)
by: Wang, Kai, et al.
Published: (2026)
$R^2$-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical Reasoning
by: Kang, Mintong, et al.
Published: (2024)
by: Kang, Mintong, et al.
Published: (2024)
RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents
by: Guo, Chengquan, et al.
Published: (2025)
by: Guo, Chengquan, et al.
Published: (2025)
BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown Attacks
by: Miao, Rui, et al.
Published: (2025)
by: Miao, Rui, et al.
Published: (2025)
CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks
by: Li, Siyuan, et al.
Published: (2026)
by: Li, Siyuan, et al.
Published: (2026)
BlueCodeAgent: A Blue Teaming Agent Enabled by Automated Red Teaming for CodeGen AI
by: Guo, Chengquan, et al.
Published: (2025)
by: Guo, Chengquan, et al.
Published: (2025)
WebGuard: Building a Generalizable Guardrail for Web Agents
by: Zheng, Boyuan, et al.
Published: (2025)
by: Zheng, Boyuan, et al.
Published: (2025)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
by: Tu, Xinming, et al.
Published: (2026)
by: Tu, Xinming, et al.
Published: (2026)
X-Guard: Multilingual Guard Agent for Content Moderation
by: Upadhayay, Bibek, et al.
Published: (2025)
by: Upadhayay, Bibek, et al.
Published: (2025)
LLM-PBE: Assessing Data Privacy in Large Language Models
by: Li, Qinbin, et al.
Published: (2024)
by: Li, Qinbin, et al.
Published: (2024)
Effective and Efficient Federated Tree Learning on Hybrid Data
by: Li, Qinbin, et al.
Published: (2023)
by: Li, Qinbin, et al.
Published: (2023)
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
AgentGuard: Runtime Verification of AI Agents
by: Koohestani, Roham
Published: (2025)
by: Koohestani, Roham
Published: (2025)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
by: Guo, Chengquan, et al.
Published: (2024)
by: Guo, Chengquan, et al.
Published: (2024)
Self-Guard: Empower the LLM to Safeguard Itself
by: Wang, Zezhong, et al.
Published: (2023)
by: Wang, Zezhong, et al.
Published: (2023)
INFA-Guard: Mitigating Malicious Propagation via Infection-Aware Safeguarding in LLM-Based Multi-Agent Systems
by: Zhou, Yijin, et al.
Published: (2026)
by: Zhou, Yijin, et al.
Published: (2026)
AgentGuard: An Attribute-Based Access Control Framework for Tool-Use LLM-Based Agent
by: Luo, Jiaqi, et al.
Published: (2026)
by: Luo, Jiaqi, et al.
Published: (2026)
MindGuard: Intrinsic Decision Inspection for Securing LLM Agents Against Metadata Poisoning
by: Wang, Zhiqiang, et al.
Published: (2025)
by: Wang, Zhiqiang, et al.
Published: (2025)
ProGuard: Towards Proactive Multimodal Safeguard
by: Yu, Shaohan, et al.
Published: (2025)
by: Yu, Shaohan, et al.
Published: (2025)
How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior
by: Xiong, Zidi, et al.
Published: (2025)
by: Xiong, Zidi, et al.
Published: (2025)
CBD: A Certified Backdoor Detector Based on Local Dominant Probability
by: Xiang, Zhen, et al.
Published: (2023)
by: Xiang, Zhen, et al.
Published: (2023)
AutoTool: Efficient Tool Selection for Large Language Model Agents
by: Jia, Jingyi, et al.
Published: (2025)
by: Jia, Jingyi, et al.
Published: (2025)
PeerGuard: Defending Multi-Agent Systems Against Backdoor Attacks Through Mutual Reasoning
by: Fan, Falong, et al.
Published: (2025)
by: Fan, Falong, et al.
Published: (2025)
GLiGuard: Schema-Conditioned Classification for LLM Safeguard
by: Zaratiana, Urchade, et al.
Published: (2026)
by: Zaratiana, Urchade, et al.
Published: (2026)
ProbGuard: Probabilistic Runtime Monitoring for LLM Agent Safety
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
by: Lin, Lixing, et al.
Published: (2026)
by: Lin, Lixing, et al.
Published: (2026)
A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
by: Wei, Qianshan, et al.
Published: (2025)
by: Wei, Qianshan, et al.
Published: (2025)
Improving Privacy-Preserving Vertical Federated Learning by Efficient Communication with ADMM
by: Xie, Chulin, et al.
Published: (2022)
by: Xie, Chulin, et al.
Published: (2022)
PerfGuard: A Performance-Aware Agent for Visual Content Generation
by: Chen, Zhipeng, et al.
Published: (2026)
by: Chen, Zhipeng, et al.
Published: (2026)
TransLinkGuard: Safeguarding Transformer Models Against Model Stealing in Edge Deployment
by: Li, Qinfeng, et al.
Published: (2024)
by: Li, Qinfeng, et al.
Published: (2024)
PropGuard: Safeguarding LLM-MAS via Propagation-Aware Exploration and Remediation
by: Yan, Bingyu, et al.
Published: (2026)
by: Yan, Bingyu, et al.
Published: (2026)
ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments
by: Wang, Yuquan, et al.
Published: (2025)
by: Wang, Yuquan, et al.
Published: (2025)
CIBER: A Comprehensive Benchmark for Security Evaluation of Code Interpreter Agents
by: Ba, Lei, et al.
Published: (2026)
by: Ba, Lei, et al.
Published: (2026)
RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents
by: Xiao, Wenjie, et al.
Published: (2026)
by: Xiao, Wenjie, et al.
Published: (2026)
VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection
by: Nie, Yuzhou, et al.
Published: (2025)
by: Nie, Yuzhou, et al.
Published: (2025)
BraveGuard: From Open-World Threats to Safer Computer-Use Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
MirrorGuard: Toward Secure Computer-Use Agents via Simulation-to-Real Reasoning Correction
by: Zhang, Wenqi, et al.
Published: (2026)
by: Zhang, Wenqi, et al.
Published: (2026)
Similar Items
-
GuardReasoner: Towards Reasoning-based LLM Safeguards
by: Liu, Yue, et al.
Published: (2025) -
AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
by: Chen, Zhaorun, et al.
Published: (2024) -
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
by: Chen, Yulin, et al.
Published: (2026) -
TrinityGuard: A Unified Framework for Safeguarding Multi-Agent Systems
by: Wang, Kai, et al.
Published: (2026) -
$R^2$-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical Reasoning
by: Kang, Mintong, et al.
Published: (2024)