SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Zhe, Ying, Zonghao, Zhang, Wenxin, Zou, Quanchen, Zhang, Deyue, Yang, Dongdong, Zhang, Xiangzheng, Peng, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs
by: Xu, Wenzhuo, et al.
Published: (2026)
by: Xu, Wenzhuo, et al.
Published: (2026)
Reasoning-Oriented Programming: Chaining Semantic Gadgets to Jailbreak Large Vision Language Models
by: Zou, Quanchen, et al.
Published: (2026)
by: Zou, Quanchen, et al.
Published: (2026)
Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application
by: Xu, Wenzhuo, et al.
Published: (2025)
by: Xu, Wenzhuo, et al.
Published: (2025)
Robust Privacy: Inference-Time Privacy through Certified Robustness
by: Jin, Jiankai, et al.
Published: (2026)
by: Jin, Jiankai, et al.
Published: (2026)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking
by: Zou, Quanchen, et al.
Published: (2025)
by: Zou, Quanchen, et al.
Published: (2025)
Two Frames Matter: A Temporal Attack for Text-to-Video Model Jailbreaking
by: Chen, Moyang, et al.
Published: (2026)
by: Chen, Moyang, et al.
Published: (2026)
Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
by: Zhang, Deyue, et al.
Published: (2025)
by: Zhang, Deyue, et al.
Published: (2025)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
by: Mu, Junjie, et al.
Published: (2025)
by: Mu, Junjie, et al.
Published: (2025)
AgentBreeder: Mitigating the AI Safety Risks of Multi-Agent Scaffolds via Self-Improvement
by: Rosser, J, et al.
Published: (2025)
by: Rosser, J, et al.
Published: (2025)
Portable Agent Memory: A Protocol for Cryptographically-Verified Memory Transfer Across Heterogeneous AI Agents
by: Ravindran, Santhosh Kumar
Published: (2026)
by: Ravindran, Santhosh Kumar
Published: (2026)
Multimodal Multi-Agent Ransomware Analysis Using AutoGen
by: Khan, Asifullah, et al.
Published: (2026)
by: Khan, Asifullah, et al.
Published: (2026)
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
by: Ying, Zonghao, et al.
Published: (2026)
by: Ying, Zonghao, et al.
Published: (2026)
Evolving Deception: When Agents Evolve, Deception Wins
by: Ying, Zonghao, et al.
Published: (2026)
by: Ying, Zonghao, et al.
Published: (2026)
BlockA2A: Towards Secure and Verifiable Agent-to-Agent Interoperability
by: Zou, Zhenhua, et al.
Published: (2025)
by: Zou, Zhenhua, et al.
Published: (2025)
SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
Demystifying and Detecting Agentic Workflow Injection Vulnerabilities in GitHub Actions
by: Wang, Shenao, et al.
Published: (2026)
by: Wang, Shenao, et al.
Published: (2026)
SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
PoseGuard: Pose-Guided Generation with Safety Guardrails
by: Wang, Kongxin, et al.
Published: (2025)
by: Wang, Kongxin, et al.
Published: (2025)
An Organization-Scoped LLM Agent Runtime Architecture for Regulated Cybersecurity Operations
by: Fatouros, George, et al.
Published: (2026)
by: Fatouros, George, et al.
Published: (2026)
Control Physiology: An Agent-Based Model of FAIR-CAM Dynamics
by: Jones, Jack, et al.
Published: (2026)
by: Jones, Jack, et al.
Published: (2026)
Tool Receipts, Not Zero-Knowledge Proofs: Practical Hallucination Detection for AI Agents
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
by: Ying, Zonghao, et al.
Published: (2024)
by: Ying, Zonghao, et al.
Published: (2024)
Defense Against Indirect Prompt Injection via Tool Result Parsing
by: Yu, Qiang, et al.
Published: (2026)
by: Yu, Qiang, et al.
Published: (2026)
When the Agent Is the Adversary: Architectural Requirements for Agentic AI Containment After the April 2026 Frontier Model Escape
by: Mitchell, Richard Joseph
Published: (2026)
by: Mitchell, Richard Joseph
Published: (2026)
MemoPhishAgent: Memory-Augmented Multi-Modal LLM Agent for Phishing URL Detection
by: Chen, Xuan, et al.
Published: (2026)
by: Chen, Xuan, et al.
Published: (2026)
Provably Secure Agent Guardrail
by: Wu, Benlong, et al.
Published: (2026)
by: Wu, Benlong, et al.
Published: (2026)
Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Execution
by: Alizadeh, Meysam, et al.
Published: (2025)
by: Alizadeh, Meysam, et al.
Published: (2025)
AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
AgentSys: Secure and Dynamic LLM Agents Through Explicit Hierarchical Memory Management
by: Wen, Ruoyao, et al.
Published: (2026)
by: Wen, Ruoyao, et al.
Published: (2026)
A V2X-based Privacy Preserving Federated Measuring and Learning System
by: Alekszejenkó, Levente, et al.
Published: (2024)
by: Alekszejenkó, Levente, et al.
Published: (2024)
Quantigence: A Multi-Agent AI Framework for Quantum Security Research
by: Alquwayfili, Abdulmalik
Published: (2025)
by: Alquwayfili, Abdulmalik
Published: (2025)
Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
by: Kang, Mintong, et al.
Published: (2025)
by: Kang, Mintong, et al.
Published: (2025)
DLP: towards active defense against backdoor attacks with decoupled learning process
by: Ying, Zonghao, et al.
Published: (2024)
by: Ying, Zonghao, et al.
Published: (2024)
NBA: defensive distillation for backdoor removal via neural behavior alignment
by: Ying, Zonghao, et al.
Published: (2024)
by: Ying, Zonghao, et al.
Published: (2024)
A Comparative Evaluation of AI Agent Security Guardrails
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
Enhancing Guardrails for Safe and Secure Healthcare AI
by: Gangavarapu, Ananya
Published: (2024)
by: Gangavarapu, Ananya
Published: (2024)
Hoist with His Own Petard: Inducing Guardrails to Facilitate Denial-of-Service Attacks on Retrieval-Augmented Generation of LLMs
by: Suo, Pan, et al.
Published: (2025)
by: Suo, Pan, et al.
Published: (2025)
Interpretable LLM Guardrails via Sparse Representation Steering
by: He, Zeqing, et al.
Published: (2025)
by: He, Zeqing, et al.
Published: (2025)
Similar Items
-
DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs
by: Xu, Wenzhuo, et al.
Published: (2026) -
Reasoning-Oriented Programming: Chaining Semantic Gadgets to Jailbreak Large Vision Language Models
by: Zou, Quanchen, et al.
Published: (2026) -
Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application
by: Xu, Wenzhuo, et al.
Published: (2025) -
Robust Privacy: Inference-Time Privacy through Certified Robustness
by: Jin, Jiankai, et al.
Published: (2026) -
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
by: Ying, Zonghao, et al.
Published: (2025)