A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Qianshan, Yang, Tengchao, Wang, Yaochen, Li, Xinfeng, Li, Lijun, Yin, Zhenfei, Zhan, Yi, Holz, Thorsten, Lin, Zhiqiang, Wang, XiaoFeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914071807262720
author Wei, Qianshan
Yang, Tengchao
Wang, Yaochen
Li, Xinfeng
Li, Lijun
Yin, Zhenfei
Zhan, Yi
Holz, Thorsten
Lin, Zhiqiang
Wang, XiaoFeng
author_facet Wei, Qianshan
Yang, Tengchao
Wang, Yaochen
Li, Xinfeng
Li, Lijun
Yin, Zhenfei
Zhan, Yi
Holz, Thorsten
Lin, Zhiqiang
Wang, XiaoFeng
contents Large Language Model (LLM) agents use memory to learn from past interactions, enabling autonomous planning and decision-making in complex environments. However, this reliance on memory introduces a critical security risk: an adversary can inject seemingly harmless records into an agent's memory to manipulate its future behavior. This vulnerability is characterized by two core aspects: First, the malicious effect of injected records is only activated within a specific context, making them hard to detect when individual memory entries are audited in isolation. Second, once triggered, the manipulation can initiate a self-reinforcing error cycle: the corrupted outcome is stored as precedent, which not only amplifies the initial error but also progressively lowers the threshold for similar attacks in the future. To address these challenges, we introduce A-MemGuard (Agent-Memory Guard), the first proactive defense framework for LLM agent memory. The core idea of our work is the insight that memory itself must become both self-checking and self-correcting. Without modifying the agent's core architecture, A-MemGuard combines two mechanisms: (1) consensus-based validation, which detects anomalies by comparing reasoning paths derived from multiple related memories and (2) a dual-memory structure, where detected failures are distilled into ``lessons'' stored separately and consulted before future actions, breaking error cycles and enabling adaptation. Comprehensive evaluations on multiple benchmarks show that A-MemGuard effectively cuts attack success rates by over 95% while incurring a minimal utility cost. This work shifts LLM memory security from static filtering to a proactive, experience-driven model where defenses strengthen over time. Our code is available in https://github.com/TangciuYueng/AMemGuard
format Preprint
id arxiv_https___arxiv_org_abs_2510_02373
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
Wei, Qianshan
Yang, Tengchao
Wang, Yaochen
Li, Xinfeng
Li, Lijun
Yin, Zhenfei
Zhan, Yi
Holz, Thorsten
Lin, Zhiqiang
Wang, XiaoFeng
Cryptography and Security
Artificial Intelligence
Large Language Model (LLM) agents use memory to learn from past interactions, enabling autonomous planning and decision-making in complex environments. However, this reliance on memory introduces a critical security risk: an adversary can inject seemingly harmless records into an agent's memory to manipulate its future behavior. This vulnerability is characterized by two core aspects: First, the malicious effect of injected records is only activated within a specific context, making them hard to detect when individual memory entries are audited in isolation. Second, once triggered, the manipulation can initiate a self-reinforcing error cycle: the corrupted outcome is stored as precedent, which not only amplifies the initial error but also progressively lowers the threshold for similar attacks in the future. To address these challenges, we introduce A-MemGuard (Agent-Memory Guard), the first proactive defense framework for LLM agent memory. The core idea of our work is the insight that memory itself must become both self-checking and self-correcting. Without modifying the agent's core architecture, A-MemGuard combines two mechanisms: (1) consensus-based validation, which detects anomalies by comparing reasoning paths derived from multiple related memories and (2) a dual-memory structure, where detected failures are distilled into ``lessons'' stored separately and consulted before future actions, breaking error cycles and enabling adaptation. Comprehensive evaluations on multiple benchmarks show that A-MemGuard effectively cuts attack success rates by over 95% while incurring a minimal utility cost. This work shifts LLM memory security from static filtering to a proactive, experience-driven model where defenses strengthen over time. Our code is available in https://github.com/TangciuYueng/AMemGuard
title A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2510.02373