Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Patlan, Atharv Singh, Sheng, Peiyao, Hebbar, S. Ashwin, Mittal, Prateek, Viswanath, Pramod |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Context manipulation attacks : Web agents are susceptible to corrupted memory
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
MURMUR: Using cross-user chatter to break collaborative language agents in groups
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
Temporal Attack Pattern Detection in Multi-Agent AI Workflows: An Open Framework for Training Trace-Based Security Models
von: Del Rosario, Ron F.
Veröffentlicht: (2025)
von: Del Rosario, Ron F.
Veröffentlicht: (2025)
Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
von: Young, Richard J.
Veröffentlicht: (2025)
von: Young, Richard J.
Veröffentlicht: (2025)
Exploiting Web Search Tools of AI Agents for Data Exfiltration
von: Rall, Dennis, et al.
Veröffentlicht: (2025)
von: Rall, Dennis, et al.
Veröffentlicht: (2025)
Hidden in Memory: Sleeper Memory Poisoning in LLM Agents
von: Pulipaka, Sidharth, et al.
Veröffentlicht: (2026)
von: Pulipaka, Sidharth, et al.
Veröffentlicht: (2026)
Efficient LLM Safety Evaluation through Multi-Agent Debate
von: Lin, Dachuan, et al.
Veröffentlicht: (2025)
von: Lin, Dachuan, et al.
Veröffentlicht: (2025)
Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems
von: Pai, Aaditya
Veröffentlicht: (2026)
von: Pai, Aaditya
Veröffentlicht: (2026)
UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation
von: Geng, Runpeng, et al.
Veröffentlicht: (2025)
von: Geng, Runpeng, et al.
Veröffentlicht: (2025)
AI Safeguards, Generative AI and the Pandora Box: AI Safety Measures to Protect Businesses and Personal Reputation
von: Kumar, Prasanna
Veröffentlicht: (2026)
von: Kumar, Prasanna
Veröffentlicht: (2026)
DWFS-Obfuscation: Dynamic Weighted Feature Selection for Robust Malware Familial Classification under Obfuscation
von: Wei, Xingyuan, et al.
Veröffentlicht: (2025)
von: Wei, Xingyuan, et al.
Veröffentlicht: (2025)
POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models
von: Shao, Yangguang, et al.
Veröffentlicht: (2025)
von: Shao, Yangguang, et al.
Veröffentlicht: (2025)
Defending against Backdoor Attacks via Module Switching
von: Li, Weijun, et al.
Veröffentlicht: (2025)
von: Li, Weijun, et al.
Veröffentlicht: (2025)
Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible
von: Zhao, Lepeng, et al.
Veröffentlicht: (2026)
von: Zhao, Lepeng, et al.
Veröffentlicht: (2026)
VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
von: Miculicich, Lesly, et al.
Veröffentlicht: (2025)
von: Miculicich, Lesly, et al.
Veröffentlicht: (2025)
Measuring Harmfulness of Computer-Using Agents
von: Tian, Aaron Xuxiang, et al.
Veröffentlicht: (2025)
von: Tian, Aaron Xuxiang, et al.
Veröffentlicht: (2025)
MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents
von: Gowda, Ishrith
Veröffentlicht: (2026)
von: Gowda, Ishrith
Veröffentlicht: (2026)
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
von: Leonesi, Matteo, et al.
Veröffentlicht: (2026)
von: Leonesi, Matteo, et al.
Veröffentlicht: (2026)
Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies
von: Nakamura, Mason, et al.
Veröffentlicht: (2025)
von: Nakamura, Mason, et al.
Veröffentlicht: (2025)
Proof of Diligence: Cryptoeconomic Security for Rollups
von: Sheng, Peiyao, et al.
Veröffentlicht: (2024)
von: Sheng, Peiyao, et al.
Veröffentlicht: (2024)
Unconditionally Safe Light Client
von: Moshrefi, Niusha, et al.
Veröffentlicht: (2024)
von: Moshrefi, Niusha, et al.
Veröffentlicht: (2024)
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
von: Zhang, Tian, et al.
Veröffentlicht: (2026)
von: Zhang, Tian, et al.
Veröffentlicht: (2026)
Reducing Information Overload: Because Even Security Experts Need to Blink
von: Kuehn, Philipp, et al.
Veröffentlicht: (2022)
von: Kuehn, Philipp, et al.
Veröffentlicht: (2022)
Training AI to be Loyal
von: Oh, Sewoong, et al.
Veröffentlicht: (2025)
von: Oh, Sewoong, et al.
Veröffentlicht: (2025)
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
von: Gu, Yongtong, et al.
Veröffentlicht: (2026)
von: Gu, Yongtong, et al.
Veröffentlicht: (2026)
From Multi-Agent Systems and the Semantic Web to Agentic AI: A Unified Narrative of the Web of Agents
von: Petrova, Tatiana, et al.
Veröffentlicht: (2025)
von: Petrova, Tatiana, et al.
Veröffentlicht: (2025)
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
von: Hill, Brennen, et al.
Veröffentlicht: (2025)
von: Hill, Brennen, et al.
Veröffentlicht: (2025)
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
von: Xin, Yuan, et al.
Veröffentlicht: (2025)
von: Xin, Yuan, et al.
Veröffentlicht: (2025)
Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems
von: Hackett, William, et al.
Veröffentlicht: (2025)
von: Hackett, William, et al.
Veröffentlicht: (2025)
Guardians of the Web: The Evolution and Future of Website Information Security
von: Islam, Md Saiful, et al.
Veröffentlicht: (2025)
von: Islam, Md Saiful, et al.
Veröffentlicht: (2025)
SecEmb: Sparsity-Aware Secure Federated Learning of On-Device Recommender System with Large Embedding
von: Mai, Peihua, et al.
Veröffentlicht: (2025)
von: Mai, Peihua, et al.
Veröffentlicht: (2025)
ConfusionPrompt: Practical Private Inference for Online Large Language Models
von: Mai, Peihua, et al.
Veröffentlicht: (2023)
von: Mai, Peihua, et al.
Veröffentlicht: (2023)
sudoLLM: On Multi-role Alignment of Language Models
von: Saha, Soumadeep, et al.
Veröffentlicht: (2025)
von: Saha, Soumadeep, et al.
Veröffentlicht: (2025)
Split-and-Denoise: Protect large language model inference with local differential privacy
von: Mai, Peihua, et al.
Veröffentlicht: (2023)
von: Mai, Peihua, et al.
Veröffentlicht: (2023)
Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
von: Adiletta, Andrew, et al.
Veröffentlicht: (2025)
von: Adiletta, Andrew, et al.
Veröffentlicht: (2025)
Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs)
von: Verma, Apurv, et al.
Veröffentlicht: (2024)
von: Verma, Apurv, et al.
Veröffentlicht: (2024)
Train to Defend: First Defense Against Cryptanalytic Neural Network Parameter Extraction Attacks
von: Kurian, Ashley, et al.
Veröffentlicht: (2025)
von: Kurian, Ashley, et al.
Veröffentlicht: (2025)
Detecting Prompt Injection Attacks Against Application Using Classifiers
von: Shaheer, Safwan, et al.
Veröffentlicht: (2025)
von: Shaheer, Safwan, et al.
Veröffentlicht: (2025)
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
von: Shaheer, Safwan, et al.
Veröffentlicht: (2025)
von: Shaheer, Safwan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Context manipulation attacks : Web agents are susceptible to corrupted memory
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025) -
MURMUR: Using cross-user chatter to break collaborative language agents in groups
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025) -
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025) -
Temporal Attack Pattern Detection in Multi-Agent AI Workflows: An Open Framework for Training Trace-Based Security Models
von: Del Rosario, Ron F.
Veröffentlicht: (2025) -
Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
von: Young, Richard J.
Veröffentlicht: (2025)