Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents
Fuente:
arXiv
Saved in:
| Main Author: | Leong, Jun Wen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attacks and Defenses Against LLM Fingerprinting
by: Kurian, Kevin, et al.
Published: (2025)
by: Kurian, Kevin, et al.
Published: (2025)
The Autonomy Tax: Defense Training Breaks LLM Agents
by: Li, Shawn, et al.
Published: (2026)
by: Li, Shawn, et al.
Published: (2026)
Seed Hijacking of LLM Sampling and Quantum Random Number Defense
by: You, Ziyang, et al.
Published: (2026)
by: You, Ziyang, et al.
Published: (2026)
Context manipulation attacks : Web agents are susceptible to corrupted memory
by: Patlan, Atharv Singh, et al.
Published: (2025)
by: Patlan, Atharv Singh, et al.
Published: (2025)
Disrupting Model Merging: A Parameter-Level Defense Without Sacrificing Accuracy
by: Junhao, Wei, et al.
Published: (2025)
by: Junhao, Wei, et al.
Published: (2025)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
by: Zhang, Heyi, et al.
Published: (2025)
by: Zhang, Heyi, et al.
Published: (2025)
Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs
by: Maiorano, Alexandre Cristovão
Published: (2026)
by: Maiorano, Alexandre Cristovão
Published: (2026)
On the use of neurosymbolic AI for defending against cyber attacks
by: Grov, Gudmund, et al.
Published: (2024)
by: Grov, Gudmund, et al.
Published: (2024)
Exploring the limits of strong membership inference attacks on large language models
by: Hayes, Jamie, et al.
Published: (2025)
by: Hayes, Jamie, et al.
Published: (2025)
Effective backdoor attack on graph neural networks in link prediction tasks
by: Dai, Jiazhu, et al.
Published: (2024)
by: Dai, Jiazhu, et al.
Published: (2024)
Enhancing web traffic attacks identification through ensemble methods and feature selection
by: Urda, Daniel, et al.
Published: (2024)
by: Urda, Daniel, et al.
Published: (2024)
Towards the generation of hierarchical attack models from cybersecurity vulnerabilities using language models
by: Sowka, Kacper, et al.
Published: (2024)
by: Sowka, Kacper, et al.
Published: (2024)
Machine Learning needs Better Randomness Standards: Randomised Smoothing and PRNG-based attacks
by: Dahiya, Pranav, et al.
Published: (2023)
by: Dahiya, Pranav, et al.
Published: (2023)
Optimal Defenses Against Gradient Reconstruction Attacks
by: Chen, Yuxiao, et al.
Published: (2024)
by: Chen, Yuxiao, et al.
Published: (2024)
Mitigating the Structural Bias in Graph Adversarial Defenses
by: Fang, Junyuan, et al.
Published: (2025)
by: Fang, Junyuan, et al.
Published: (2025)
Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses
by: Pawlak, Stanisław, et al.
Published: (2025)
by: Pawlak, Stanisław, et al.
Published: (2025)
A Causal Perspective for Enhancing Jailbreak Attack and Defense
by: Pan, Licheng, et al.
Published: (2026)
by: Pan, Licheng, et al.
Published: (2026)
Stealthy Poisoning Attacks Bypass Defenses in Regression Settings
by: Carnerero-Cano, Javier, et al.
Published: (2026)
by: Carnerero-Cano, Javier, et al.
Published: (2026)
SHIELD: Secure Hypernetworks for Incremental Expansion Learning Defense
by: Krukowski, Patryk, et al.
Published: (2025)
by: Krukowski, Patryk, et al.
Published: (2025)
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
by: Min, Rui, et al.
Published: (2024)
by: Min, Rui, et al.
Published: (2024)
An Investigation into the Performances of the State-of-the-art Machine Learning Approaches for Various Cyber-attack Detection: A Survey
by: Ige, Tosin, et al.
Published: (2024)
by: Ige, Tosin, et al.
Published: (2024)
Can Adversarial Code Comments Fool AI Security Reviewers -- Large-Scale Empirical Study of Comment-Based Attacks and Defenses Against LLM Code Analysis
by: Thornton, Scott
Published: (2026)
by: Thornton, Scott
Published: (2026)
Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques
by: Jaffal, Niveen O., et al.
Published: (2025)
by: Jaffal, Niveen O., et al.
Published: (2025)
Attacks and Defenses for Generative Diffusion Models: A Comprehensive Survey
by: Truong, Vu Tuan, et al.
Published: (2024)
by: Truong, Vu Tuan, et al.
Published: (2024)
Training RL Agents for Multi-Objective Network Defense Tasks
by: Molina-Markham, Andres, et al.
Published: (2025)
by: Molina-Markham, Andres, et al.
Published: (2025)
Combining Stochastic Defenses to Resist Gradient Inversion: An Ablation Study
by: Scheliga, Daniel, et al.
Published: (2022)
by: Scheliga, Daniel, et al.
Published: (2022)
Accuracy of TextFooler black box adversarial attacks on 01 loss sign activation neural network ensemble
by: Xue, Yunzhe, et al.
Published: (2024)
by: Xue, Yunzhe, et al.
Published: (2024)
RAB$^2$-DEF: Dynamic and explainable defense against adversarial attacks in Federated Learning to fair poor clients
by: Rodríguez-Barroso, Nuria, et al.
Published: (2024)
by: Rodríguez-Barroso, Nuria, et al.
Published: (2024)
Semantic Chameleon: Corpus-Dependent Poisoning Attacks and Defenses in RAG Systems
by: Thornton, Scott
Published: (2026)
by: Thornton, Scott
Published: (2026)
A Survey on Model Extraction Attacks and Defenses for Large Language Models
by: Zhao, Kaixiang, et al.
Published: (2025)
by: Zhao, Kaixiang, et al.
Published: (2025)
Optimizing Cyber Defense in Dynamic Active Directories through Reinforcement Learning
by: Goel, Diksha, et al.
Published: (2024)
by: Goel, Diksha, et al.
Published: (2024)
A Survey of Model Extraction Attacks and Defenses in Distributed Computing Environments
by: Zhao, Kaixiang, et al.
Published: (2025)
by: Zhao, Kaixiang, et al.
Published: (2025)
Adversarial Reinforcement Learning for Offensive and Defensive Agents in a Simulated Zero-Sum Network Environment
by: Shahid, Abrar, et al.
Published: (2025)
by: Shahid, Abrar, et al.
Published: (2025)
IRSKG: Unified Intrusion Response System Knowledge Graph Ontology for Cyber Defense
by: Panigrahi, Damodar, et al.
Published: (2024)
by: Panigrahi, Damodar, et al.
Published: (2024)
Defending the Edge: Representative-Attention Defense against Backdoor Attacks in Federated Learning
by: Obioma, Chibueze Peace, et al.
Published: (2025)
by: Obioma, Chibueze Peace, et al.
Published: (2025)
Adversarial Robustness in Financial Machine Learning: Defenses, Economic Impact, and Governance Evidence
by: Baviskar, Samruddhi
Published: (2025)
by: Baviskar, Samruddhi
Published: (2025)
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
A Systematic Survey of Model Extraction Attacks and Defenses: State-of-the-Art and Perspectives
by: Zhao, Kaixiang, et al.
Published: (2025)
by: Zhao, Kaixiang, et al.
Published: (2025)
Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
by: Konrad, Phongsakon Mark, et al.
Published: (2026)
by: Konrad, Phongsakon Mark, et al.
Published: (2026)
DeepStage: Learning Autonomous Defense Policies Against Multi-Stage APT Campaigns
by: Phan, Trung V., et al.
Published: (2026)
by: Phan, Trung V., et al.
Published: (2026)
Similar Items
-
Attacks and Defenses Against LLM Fingerprinting
by: Kurian, Kevin, et al.
Published: (2025) -
The Autonomy Tax: Defense Training Breaks LLM Agents
by: Li, Shawn, et al.
Published: (2026) -
Seed Hijacking of LLM Sampling and Quantum Random Number Defense
by: You, Ziyang, et al.
Published: (2026) -
Context manipulation attacks : Web agents are susceptible to corrupted memory
by: Patlan, Atharv Singh, et al.
Published: (2025) -
Disrupting Model Merging: A Parameter-Level Defense Without Sacrificing Accuracy
by: Junhao, Wei, et al.
Published: (2025)