Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Boyang, Tan, Yicong, Shen, Yun, Salem, Ahmed, Backes, Michael, Zannettou, Savvas, Zhang, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
von: Shen, Xinyue, et al.
Veröffentlicht: (2025)
von: Shen, Xinyue, et al.
Veröffentlicht: (2025)
Vera Verto: Multimodal Hijacking Attack
von: Zhang, Minxing, et al.
Veröffentlicht: (2024)
von: Zhang, Minxing, et al.
Veröffentlicht: (2024)
Assessing Deanonymization Risks with Stylometry-Assisted LLM Agent
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTs
von: Shen, Xinyue, et al.
Veröffentlicht: (2025)
von: Shen, Xinyue, et al.
Veröffentlicht: (2025)
Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks
von: Chu, Junjie, et al.
Veröffentlicht: (2026)
von: Chu, Junjie, et al.
Veröffentlicht: (2026)
Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution
von: Wu, Yixin, et al.
Veröffentlicht: (2024)
von: Wu, Yixin, et al.
Veröffentlicht: (2024)
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
Synthetic Artifact Auditing: Tracing LLM-Generated Synthetic Data Usage in Downstream Applications
von: Wu, Yixin, et al.
Veröffentlicht: (2025)
von: Wu, Yixin, et al.
Veröffentlicht: (2025)
Compromising Embodied Agents with Contextual Backdoor Attacks
von: Liu, Aishan, et al.
Veröffentlicht: (2024)
von: Liu, Aishan, et al.
Veröffentlicht: (2024)
The Challenge of Identifying the Origin of Black-Box Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2025)
von: Yang, Ziqing, et al.
Veröffentlicht: (2025)
Voice Jailbreak Attacks Against GPT-4o
von: Shen, Xinyue, et al.
Veröffentlicht: (2024)
von: Shen, Xinyue, et al.
Veröffentlicht: (2024)
Prompt Stealing Attacks Against Text-to-Image Generation Models
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
von: Srivastava, Saksham Sahai, et al.
Veröffentlicht: (2025)
von: Srivastava, Saksham Sahai, et al.
Veröffentlicht: (2025)
HARP: Measuring Harm Amplification in Multi-Agent LLM Systems
von: Rahman, Md Hafizur, et al.
Veröffentlicht: (2026)
von: Rahman, Md Hafizur, et al.
Veröffentlicht: (2026)
Instruction Backdoor Attacks Against Customized LLMs
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
Understanding Data Importance in Machine Learning Attacks: Does Valuable Data Pose Greater Harm?
von: Wen, Rui, et al.
Veröffentlicht: (2024)
von: Wen, Rui, et al.
Veröffentlicht: (2024)
Composite Backdoor Attacks Against Large Language Models
von: Huang, Hai, et al.
Veröffentlicht: (2023)
von: Huang, Hai, et al.
Veröffentlicht: (2023)
MGTBench: Benchmarking Machine-Generated Text Detection
von: He, Xinlei, et al.
Veröffentlicht: (2023)
von: He, Xinlei, et al.
Veröffentlicht: (2023)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
The Autonomy Tax: Defense Training Breaks LLM Agents
von: Li, Shawn, et al.
Veröffentlicht: (2026)
von: Li, Shawn, et al.
Veröffentlicht: (2026)
CTFusion: A CTF-based Benchmark for LLM Agent Evaluation
von: Lee, Dongjun, et al.
Veröffentlicht: (2026)
von: Lee, Dongjun, et al.
Veröffentlicht: (2026)
Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions
von: Qu, Yiting, et al.
Veröffentlicht: (2025)
von: Qu, Yiting, et al.
Veröffentlicht: (2025)
SoK: Data Reconstruction Attacks Against Machine Learning Models: Definition, Metrics, and Benchmark
von: Wen, Rui, et al.
Veröffentlicht: (2025)
von: Wen, Rui, et al.
Veröffentlicht: (2025)
Transferable Availability Poisoning Attacks
von: Liu, Yiyong, et al.
Veröffentlicht: (2023)
von: Liu, Yiyong, et al.
Veröffentlicht: (2023)
Agent-ToM: Learning to Monitor Autonomous LLM Agents via Theory-of-Mind Reasoning
von: Ahmed, Nesreen K., et al.
Veröffentlicht: (2026)
von: Ahmed, Nesreen K., et al.
Veröffentlicht: (2026)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2024)
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2024)
Excessive Reasoning Attack on Reasoning LLMs
von: Si, Wai Man, et al.
Veröffentlicht: (2025)
von: Si, Wai Man, et al.
Veröffentlicht: (2025)
UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images
von: Qu, Yiting, et al.
Veröffentlicht: (2024)
von: Qu, Yiting, et al.
Veröffentlicht: (2024)
SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
von: Lee, Hwiwon, et al.
Veröffentlicht: (2025)
von: Lee, Hwiwon, et al.
Veröffentlicht: (2025)
Evaluating Generalization Mechanisms in Autonomous Cyber Attack Agents
von: Lukáš, Ondřej, et al.
Veröffentlicht: (2026)
von: Lukáš, Ondřej, et al.
Veröffentlicht: (2026)
Provably Cost-Sensitive Adversarial Defense via Randomized Smoothing
von: Xin, Yuan, et al.
Veröffentlicht: (2023)
von: Xin, Yuan, et al.
Veröffentlicht: (2023)
Privacy Amplification Through Synthetic Data: Insights from Linear Regression
von: Pierquin, Clément, et al.
Veröffentlicht: (2025)
von: Pierquin, Clément, et al.
Veröffentlicht: (2025)
AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
von: Wu, Yixin, et al.
Veröffentlicht: (2025)
von: Wu, Yixin, et al.
Veröffentlicht: (2025)
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
von: Bazinska, Julia, et al.
Veröffentlicht: (2025)
von: Bazinska, Julia, et al.
Veröffentlicht: (2025)
Inference-Time Backdoors via Chat Templates: From LLM Supply Chains to Agentic System Compromise
von: Fogel, Ariel, et al.
Veröffentlicht: (2026)
von: Fogel, Ariel, et al.
Veröffentlicht: (2026)
Memory-Induced Tool-Drift in LLM Agents
von: Dabas, Mahavir, et al.
Veröffentlicht: (2026)
von: Dabas, Mahavir, et al.
Veröffentlicht: (2026)
Mitigating Error Amplification in Fast Adversarial Training
von: Zhao, Mengnan, et al.
Veröffentlicht: (2026)
von: Zhao, Mengnan, et al.
Veröffentlicht: (2026)
Unified Mechanism-Specific Amplification by Subsampling and Group Privacy Amplification
von: Schuchardt, Jan, et al.
Veröffentlicht: (2024)
von: Schuchardt, Jan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
von: Shen, Xinyue, et al.
Veröffentlicht: (2025) -
Vera Verto: Multimodal Hijacking Attack
von: Zhang, Minxing, et al.
Veröffentlicht: (2024) -
Assessing Deanonymization Risks with Stylometry-Assisted LLM Agent
von: Zhang, Boyang, et al.
Veröffentlicht: (2026) -
SOS! Soft Prompt Attack Against Open-Source Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2024) -
GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)