SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Zhe, Ying, Zonghao, Zhang, Wenxin, Zou, Quanchen, Zhang, Deyue, Yang, Dongdong, Zhang, Xiangzheng, Peng, Hao
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914589710483456
author Liu, Zhe
Ying, Zonghao
Zhang, Wenxin
Zou, Quanchen
Zhang, Deyue
Yang, Dongdong
Zhang, Xiangzheng
Peng, Hao
author_facet Liu, Zhe
Ying, Zonghao
Zhang, Wenxin
Zou, Quanchen
Zhang, Deyue
Yang, Dongdong
Zhang, Xiangzheng
Peng, Hao
contents Recent advances in foundation models have transformed LLMs from passive conversational systems into autonomous agents capable of reasoning and tool execution. While these capabilities unlock substantial practical value, they also introduce new security risks, as adversaries can manipulate agents into performing harmful actions in real-world environments. Existing defense strategies mitigate such threats but frequently struggle to balance safety and utility, resulting in over-refusal of benign user requests. To mitigate this trade-off, we propose SafeHarbor, a novel framework designed to establish precise decision boundaries for LLM agents. Unlike static guidelines, SafeHarbor extracts context-aware defense rules through enhanced adversarial generation. We design a local hierarchical memory system for dynamic rule injection, offering a training-free, efficient, and plug-and-play solution. Furthermore, we introduce an information entropy-based self-evolution mechanism that continuously optimizes the memory structure through dynamic node splitting and merging. Extensive experiments demonstrate that SafeHarbor achieves state-of-the-art performance on both ambiguous benign tasks and explicit malicious attacks, notably attaining a peak benign utility of 63.6\% on GPT-4o while maintaining a robust refusal rate exceeding 93\% against harmful requests. The source code is publicly available at https://github.com/ljj-cyber/SafeHarbor.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05704
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
Liu, Zhe
Ying, Zonghao
Zhang, Wenxin
Zou, Quanchen
Zhang, Deyue
Yang, Dongdong
Zhang, Xiangzheng
Peng, Hao
Cryptography and Security
Artificial Intelligence
68T42
Recent advances in foundation models have transformed LLMs from passive conversational systems into autonomous agents capable of reasoning and tool execution. While these capabilities unlock substantial practical value, they also introduce new security risks, as adversaries can manipulate agents into performing harmful actions in real-world environments. Existing defense strategies mitigate such threats but frequently struggle to balance safety and utility, resulting in over-refusal of benign user requests. To mitigate this trade-off, we propose SafeHarbor, a novel framework designed to establish precise decision boundaries for LLM agents. Unlike static guidelines, SafeHarbor extracts context-aware defense rules through enhanced adversarial generation. We design a local hierarchical memory system for dynamic rule injection, offering a training-free, efficient, and plug-and-play solution. Furthermore, we introduce an information entropy-based self-evolution mechanism that continuously optimizes the memory structure through dynamic node splitting and merging. Extensive experiments demonstrate that SafeHarbor achieves state-of-the-art performance on both ambiguous benign tasks and explicit malicious attacks, notably attaining a peak benign utility of 63.6\% on GPT-4o while maintaining a robust refusal rate exceeding 93\% against harmful requests. The source code is publicly available at https://github.com/ljj-cyber/SafeHarbor.
title SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
topic Cryptography and Security
Artificial Intelligence
68T42
url https://arxiv.org/abs/2605.05704