MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Gowda, Ishrith |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hidden in Memory: Sleeper Memory Poisoning in LLM Agents
von: Pulipaka, Sidharth, et al.
Veröffentlicht: (2026)
von: Pulipaka, Sidharth, et al.
Veröffentlicht: (2026)
Binary-30K: A Heterogeneous Dataset for Deep Learning in Binary Analysis and Malware Detection
von: Bommarito II, Michael J.
Veröffentlicht: (2025)
von: Bommarito II, Michael J.
Veröffentlicht: (2025)
Attacking interpretable NLP systems
von: Abdukhamidov, Eldor, et al.
Veröffentlicht: (2025)
von: Abdukhamidov, Eldor, et al.
Veröffentlicht: (2025)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
von: Othman, Refat
Veröffentlicht: (2026)
von: Othman, Refat
Veröffentlicht: (2026)
Detecting Prompt Injection Attacks Against Application Using Classifiers
von: Shaheer, Safwan, et al.
Veröffentlicht: (2025)
von: Shaheer, Safwan, et al.
Veröffentlicht: (2025)
AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
von: Nathanson, Samuel, et al.
Veröffentlicht: (2025)
von: Nathanson, Samuel, et al.
Veröffentlicht: (2025)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
von: Young, Richard J., et al.
Veröffentlicht: (2026)
von: Young, Richard J., et al.
Veröffentlicht: (2026)
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
von: Shaheer, Safwan, et al.
Veröffentlicht: (2025)
von: Shaheer, Safwan, et al.
Veröffentlicht: (2025)
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
von: Zhang, Tian, et al.
Veröffentlicht: (2026)
von: Zhang, Tian, et al.
Veröffentlicht: (2026)
Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice
von: Ge, Yuxu
Veröffentlicht: (2026)
von: Ge, Yuxu
Veröffentlicht: (2026)
Multilingual AI-Driven Password Strength Estimation with Similarity-Based Detection
von: Palaniappan, Nikitha M., et al.
Veröffentlicht: (2026)
von: Palaniappan, Nikitha M., et al.
Veröffentlicht: (2026)
AegisShield: Democratizing Cyber Threat Modeling with Generative AI
von: Grofsky, Matthew
Veröffentlicht: (2025)
von: Grofsky, Matthew
Veröffentlicht: (2025)
Evaluating the Reliability of Digital Forensic Evidence Discovered by Large Language Model: A Case Study
von: Khatiwala, Jeel Piyushkumar, et al.
Veröffentlicht: (2026)
von: Khatiwala, Jeel Piyushkumar, et al.
Veröffentlicht: (2026)
Towards Agentic Investigation of Security Alerts
von: Eilertsen, Even, et al.
Veröffentlicht: (2026)
von: Eilertsen, Even, et al.
Veröffentlicht: (2026)
Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies
von: Cotti, Luca, et al.
Veröffentlicht: (2025)
von: Cotti, Luca, et al.
Veröffentlicht: (2025)
Privately Fine-Tuned LLMs Preserve Temporal Dynamics in Tabular Data
von: Rosenblatt, Lucas, et al.
Veröffentlicht: (2026)
von: Rosenblatt, Lucas, et al.
Veröffentlicht: (2026)
Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents
von: Sidik, Bronislav, et al.
Veröffentlicht: (2026)
von: Sidik, Bronislav, et al.
Veröffentlicht: (2026)
Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity
von: Williams-King, David, et al.
Veröffentlicht: (2025)
von: Williams-King, David, et al.
Veröffentlicht: (2025)
Quantifying Return on Security Controls in LLM Systems
von: Moulton, Richard Helder, et al.
Veröffentlicht: (2025)
von: Moulton, Richard Helder, et al.
Veröffentlicht: (2025)
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
von: Leonesi, Matteo, et al.
Veröffentlicht: (2026)
von: Leonesi, Matteo, et al.
Veröffentlicht: (2026)
Security Considerations for Multi-agent Systems
von: Nguyen, Tam, et al.
Veröffentlicht: (2026)
von: Nguyen, Tam, et al.
Veröffentlicht: (2026)
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
von: Dawson, Ads, et al.
Veröffentlicht: (2025)
von: Dawson, Ads, et al.
Veröffentlicht: (2025)
The Automation Advantage in AI Red Teaming
von: Mulla, Rob, et al.
Veröffentlicht: (2025)
von: Mulla, Rob, et al.
Veröffentlicht: (2025)
SAND: A Self-supervised and Adaptive NAS-Driven Framework for Hardware Trojan Detection
von: Pan, Zhixin, et al.
Veröffentlicht: (2025)
von: Pan, Zhixin, et al.
Veröffentlicht: (2025)
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
von: Bercovich, Ivan, et al.
Veröffentlicht: (2026)
von: Bercovich, Ivan, et al.
Veröffentlicht: (2026)
Poison in the Well: Feature Embedding Disruption in Backdoor Attacks
von: Feng, Zhou, et al.
Veröffentlicht: (2025)
von: Feng, Zhou, et al.
Veröffentlicht: (2025)
Evaluating Query Efficiency and Accuracy of Transfer Learning-based Model Extraction Attack in Federated Learning
von: Ahamed, Sayyed Farid, et al.
Veröffentlicht: (2025)
von: Ahamed, Sayyed Farid, et al.
Veröffentlicht: (2025)
David vs. Goliath: Verifiable Agent-to-Agent Jailbreaking via Reinforcement Learning
von: Nellessen, Samuel, et al.
Veröffentlicht: (2026)
von: Nellessen, Samuel, et al.
Veröffentlicht: (2026)
A Self-Improving Architecture for Dynamic Safety in Large Language Models
von: Slater, Tyler
Veröffentlicht: (2025)
von: Slater, Tyler
Veröffentlicht: (2025)
Autonomous Penetration Testing: Solving Capture-the-Flag Challenges with LLMs
von: Bakker, Isabelle, et al.
Veröffentlicht: (2025)
von: Bakker, Isabelle, et al.
Veröffentlicht: (2025)
Can AI Keep a Secret? Contextual Integrity Verification: A Provable Security Architecture for LLMs
von: Gupta, Aayush
Veröffentlicht: (2025)
von: Gupta, Aayush
Veröffentlicht: (2025)
Unlearning at Scale: Implementing the Right to be Forgotten in Large Language Models
von: X, Abdullah
Veröffentlicht: (2025)
von: X, Abdullah
Veröffentlicht: (2025)
SALLIE: Safeguarding Against Latent Language & Image Exploits
von: Azov, Guy, et al.
Veröffentlicht: (2026)
von: Azov, Guy, et al.
Veröffentlicht: (2026)
Provable Repair of Deep Neural Network Defects by Preimage Synthesis and Property Refinement
von: Ma, Jianan, et al.
Veröffentlicht: (2025)
von: Ma, Jianan, et al.
Veröffentlicht: (2025)
Retrieval Augmented Classification for Confidential Documents
von: Chang, Yeseul E., et al.
Veröffentlicht: (2026)
von: Chang, Yeseul E., et al.
Veröffentlicht: (2026)
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
von: Chakraborty, Amit, et al.
Veröffentlicht: (2025)
von: Chakraborty, Amit, et al.
Veröffentlicht: (2025)
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
von: DeLeeuw, Caleb
Veröffentlicht: (2026)
von: DeLeeuw, Caleb
Veröffentlicht: (2026)
PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models
von: Seddik, Issam, et al.
Veröffentlicht: (2025)
von: Seddik, Issam, et al.
Veröffentlicht: (2025)
Continuous Discovery of Vulnerabilities in LLM Serving Systems with Fuzzing
von: Zhao, Yunze, et al.
Veröffentlicht: (2026)
von: Zhao, Yunze, et al.
Veröffentlicht: (2026)
Teams of LLM Agents can Exploit Zero-Day Vulnerabilities
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Hidden in Memory: Sleeper Memory Poisoning in LLM Agents
von: Pulipaka, Sidharth, et al.
Veröffentlicht: (2026) -
Binary-30K: A Heterogeneous Dataset for Deep Learning in Binary Analysis and Malware Detection
von: Bommarito II, Michael J.
Veröffentlicht: (2025) -
Attacking interpretable NLP systems
von: Abdukhamidov, Eldor, et al.
Veröffentlicht: (2025) -
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
von: Othman, Refat
Veröffentlicht: (2026) -
Detecting Prompt Injection Attacks Against Application Using Classifiers
von: Shaheer, Safwan, et al.
Veröffentlicht: (2025)