Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jha, Rishi, Triedman, Harold, Bhattacharya, Arkaprabha, Shmatikov, Vitaly |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Agent Systems Execute Arbitrary Malicious Code
von: Triedman, Harold, et al.
Veröffentlicht: (2025)
von: Triedman, Harold, et al.
Veröffentlicht: (2025)
Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems
von: Jha, Rishi, et al.
Veröffentlicht: (2025)
von: Jha, Rishi, et al.
Veröffentlicht: (2025)
Deep-Research Agents Can Be Poisoned via User-Generated Content
von: Zhang, Tingwei, et al.
Veröffentlicht: (2026)
von: Zhang, Tingwei, et al.
Veröffentlicht: (2026)
Universal Zero-shot Embedding Inversion
von: Zhang, Collin, et al.
Veröffentlicht: (2025)
von: Zhang, Collin, et al.
Veröffentlicht: (2025)
Adversarial Decoding: Generating Readable Documents for Adversarial Objectives
von: Zhang, Collin, et al.
Veröffentlicht: (2024)
von: Zhang, Collin, et al.
Veröffentlicht: (2024)
Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents
von: Shafran, Avital, et al.
Veröffentlicht: (2024)
von: Shafran, Avital, et al.
Veröffentlicht: (2024)
MillStone: How Open-Minded Are LLMs?
von: Triedman, Harold, et al.
Veröffentlicht: (2025)
von: Triedman, Harold, et al.
Veröffentlicht: (2025)
Adversarial Illusions in Multi-Modal Embeddings
von: Zhang, Tingwei, et al.
Veröffentlicht: (2023)
von: Zhang, Tingwei, et al.
Veröffentlicht: (2023)
Adversarial Hubness in Multi-Modal Retrieval
von: Zhang, Tingwei, et al.
Veröffentlicht: (2024)
von: Zhang, Tingwei, et al.
Veröffentlicht: (2024)
Differential Degradation Vulnerabilities in Censorship Circumvention Systems
von: Sun, Zhen, et al.
Veröffentlicht: (2024)
von: Sun, Zhen, et al.
Veröffentlicht: (2024)
How to Steal Reasoning Without Reasoning Traces
von: Zhang, Tingwei, et al.
Veröffentlicht: (2026)
von: Zhang, Tingwei, et al.
Veröffentlicht: (2026)
Learned, Lagged, LLM-splained: LLM Responses to End User Security Questions
von: Prakash, Vijay, et al.
Veröffentlicht: (2024)
von: Prakash, Vijay, et al.
Veröffentlicht: (2024)
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
von: Xu, Chejian, et al.
Veröffentlicht: (2024)
von: Xu, Chejian, et al.
Veröffentlicht: (2024)
AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents
von: Rassul, Yassin H., et al.
Veröffentlicht: (2026)
von: Rassul, Yassin H., et al.
Veröffentlicht: (2026)
Watermarking LLM Agent Trajectories
von: Meng, Wenlong, et al.
Veröffentlicht: (2026)
von: Meng, Wenlong, et al.
Veröffentlicht: (2026)
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
von: Zhan, Qiusi, et al.
Veröffentlicht: (2024)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2024)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
von: Gomaa, Amr, et al.
Veröffentlicht: (2025)
von: Gomaa, Amr, et al.
Veröffentlicht: (2025)
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents
von: Fu, Wenjie, et al.
Veröffentlicht: (2026)
von: Fu, Wenjie, et al.
Veröffentlicht: (2026)
Multi-Agent Collaboration in Incident Response with Large Language Models
von: Liu, Zefang
Veröffentlicht: (2024)
von: Liu, Zefang
Veröffentlicht: (2024)
A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness
von: Luo, Xuan, et al.
Veröffentlicht: (2025)
von: Luo, Xuan, et al.
Veröffentlicht: (2025)
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
von: Zhang, Chiyu, et al.
Veröffentlicht: (2026)
von: Zhang, Chiyu, et al.
Veröffentlicht: (2026)
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
von: Wang, Peiran, et al.
Veröffentlicht: (2026)
von: Wang, Peiran, et al.
Veröffentlicht: (2026)
Overthinking Loops in Agents: A Structural Risk via MCP Tools
von: Lee, Yohan, et al.
Veröffentlicht: (2026)
von: Lee, Yohan, et al.
Veröffentlicht: (2026)
LeechHijack: Covert Computational Resource Exploitation in Intelligent Agent Systems
von: Zhang, Yuanhe, et al.
Veröffentlicht: (2025)
von: Zhang, Yuanhe, et al.
Veröffentlicht: (2025)
Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents
von: Chen, Yining, et al.
Veröffentlicht: (2026)
von: Chen, Yining, et al.
Veröffentlicht: (2026)
BraveGuard: From Open-World Threats to Safer Computer-Use Agents
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities
von: Ren, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Ren, Xiaoxue, et al.
Veröffentlicht: (2025)
Personalized Attacks of Social Engineering in Multi-turn Conversations: LLM Agents for Simulation and Detection
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2025)
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2025)
Rerouting LLM Routers
von: Shafran, Avital, et al.
Veröffentlicht: (2025)
von: Shafran, Avital, et al.
Veröffentlicht: (2025)
Dissecting Adversarial Robustness of Multimodal LM Agents
von: Wu, Chen Henry, et al.
Veröffentlicht: (2024)
von: Wu, Chen Henry, et al.
Veröffentlicht: (2024)
MRJ-Agent: An Effective Jailbreak Agent for Multi-Round Dialogue
von: Wang, Fengxiang, et al.
Veröffentlicht: (2024)
von: Wang, Fengxiang, et al.
Veröffentlicht: (2024)
AirGapAgent: Protecting Privacy-Conscious Conversational Agents
von: Bagdasarian, Eugene, et al.
Veröffentlicht: (2024)
von: Bagdasarian, Eugene, et al.
Veröffentlicht: (2024)
MCPShield: A Security Cognition Layer for Adaptive Trust Calibration in Model Context Protocol Agents
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2026)
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2026)
TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent
von: Meier, Dominik, et al.
Veröffentlicht: (2025)
von: Meier, Dominik, et al.
Veröffentlicht: (2025)
CoTGuard: Using Chain-of-Thought Triggering for Copyright Protection in Multi-Agent LLM Systems
von: Wen, Yan, et al.
Veröffentlicht: (2025)
von: Wen, Yan, et al.
Veröffentlicht: (2025)
AutoBnB-RAG: Enhancing Multi-Agent Incident Response with Retrieval-Augmented Generation
von: Liu, Zefang, et al.
Veröffentlicht: (2025)
von: Liu, Zefang, et al.
Veröffentlicht: (2025)
Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning
von: Hui, Zheng, et al.
Veröffentlicht: (2025)
von: Hui, Zheng, et al.
Veröffentlicht: (2025)
MAPS: A Multilingual Benchmark for Agent Performance and Security
von: Hofman, Omer, et al.
Veröffentlicht: (2025)
von: Hofman, Omer, et al.
Veröffentlicht: (2025)
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Multi-Agent Systems Execute Arbitrary Malicious Code
von: Triedman, Harold, et al.
Veröffentlicht: (2025) -
Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems
von: Jha, Rishi, et al.
Veröffentlicht: (2025) -
Deep-Research Agents Can Be Poisoned via User-Generated Content
von: Zhang, Tingwei, et al.
Veröffentlicht: (2026) -
Universal Zero-shot Embedding Inversion
von: Zhang, Collin, et al.
Veröffentlicht: (2025) -
Adversarial Decoding: Generating Readable Documents for Adversarial Objectives
von: Zhang, Collin, et al.
Veröffentlicht: (2024)