Overthinking Loops in Agents: A Structural Risk via MCP Tools
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Yohan, Jang, Jisoo, Choi, Seoyeon, Kim, Sangyeop, Choi, Seungtaek |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs
von: Kim, Sangyeop, et al.
Veröffentlicht: (2025)
von: Kim, Sangyeop, et al.
Veröffentlicht: (2025)
Can LLM Infer Risk Information From MCP Server System Logs?
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
von: Lee, Hyomin, et al.
Veröffentlicht: (2026)
von: Lee, Hyomin, et al.
Veröffentlicht: (2026)
LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops
von: Fu, Jiyuan, et al.
Veröffentlicht: (2025)
von: Fu, Jiyuan, et al.
Veröffentlicht: (2025)
AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents
von: Rassul, Yassin H., et al.
Veröffentlicht: (2026)
von: Rassul, Yassin H., et al.
Veröffentlicht: (2026)
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
von: Zhan, Qiusi, et al.
Veröffentlicht: (2024)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2024)
Q-FAKER: Query-free Hard Black-box Attack via Controlled Generation
von: Na, CheolWon, et al.
Veröffentlicht: (2025)
von: Na, CheolWon, et al.
Veröffentlicht: (2025)
STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio
von: Park, Seong-Gyu, et al.
Veröffentlicht: (2026)
von: Park, Seong-Gyu, et al.
Veröffentlicht: (2026)
Searching for Privacy Risks in LLM Agents via Simulation
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2025)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2025)
Ghost Tool Calls: Issue-Time Privacy for Speculative Agent Tools
von: Mohammadi, Bardia, et al.
Veröffentlicht: (2026)
von: Mohammadi, Bardia, et al.
Veröffentlicht: (2026)
Enabling Efficient Attack Investigation via Human-in-the-Loop Security Analysis
von: Tsegai, Saimon Amanuel, et al.
Veröffentlicht: (2022)
von: Tsegai, Saimon Amanuel, et al.
Veröffentlicht: (2022)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2023)
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2023)
Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2026)
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2026)
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
von: Mangaokar, Neal, et al.
Veröffentlicht: (2024)
von: Mangaokar, Neal, et al.
Veröffentlicht: (2024)
Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
von: Qiao, Yuxuan, et al.
Veröffentlicht: (2025)
von: Qiao, Yuxuan, et al.
Veröffentlicht: (2025)
Preventing Jailbreak Prompts as Malicious Tools for Cybercriminals: A Cyber Defense Perspective
von: Tshimula, Jean Marie, et al.
Veröffentlicht: (2024)
von: Tshimula, Jean Marie, et al.
Veröffentlicht: (2024)
Security Attacks on LLM-based Code Completion Tools
von: Cheng, Wen, et al.
Veröffentlicht: (2024)
von: Cheng, Wen, et al.
Veröffentlicht: (2024)
ContextLeak: Auditing Leakage in Private In-Context Learning Methods
von: Choi, Jacob, et al.
Veröffentlicht: (2025)
von: Choi, Jacob, et al.
Veröffentlicht: (2025)
Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning
von: Hui, Zheng, et al.
Veröffentlicht: (2025)
von: Hui, Zheng, et al.
Veröffentlicht: (2025)
Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents
von: Jha, Rishi, et al.
Veröffentlicht: (2026)
von: Jha, Rishi, et al.
Veröffentlicht: (2026)
GraphSteal: Structural Knowledge Stealing from Graph RAG via Traversal Reconstruction
von: Gu, Jinze, et al.
Veröffentlicht: (2026)
von: Gu, Jinze, et al.
Veröffentlicht: (2026)
SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
von: Xu, Chejian, et al.
Veröffentlicht: (2024)
von: Xu, Chejian, et al.
Veröffentlicht: (2024)
Assessing Deanonymization Risks with Stylometry-Assisted LLM Agent
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation
von: Li, Xingyu, et al.
Veröffentlicht: (2025)
von: Li, Xingyu, et al.
Veröffentlicht: (2025)
IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents
von: An, Hengyu, et al.
Veröffentlicht: (2025)
von: An, Hengyu, et al.
Veröffentlicht: (2025)
Watermarking LLM Agent Trajectories
von: Meng, Wenlong, et al.
Veröffentlicht: (2026)
von: Meng, Wenlong, et al.
Veröffentlicht: (2026)
Representation Bending for Large Language Model Safety
von: Yousefpour, Ashkan, et al.
Veröffentlicht: (2025)
von: Yousefpour, Ashkan, et al.
Veröffentlicht: (2025)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
von: Chae, Kyubyung, et al.
Veröffentlicht: (2025)
von: Chae, Kyubyung, et al.
Veröffentlicht: (2025)
SoK: Privacy Risks and Mitigations in Retrieval-Augmented Generation Systems
von: Bodea, Andreea-Elena, et al.
Veröffentlicht: (2026)
von: Bodea, Andreea-Elena, et al.
Veröffentlicht: (2026)
MCPShield: A Security Cognition Layer for Adaptive Trust Calibration in Model Context Protocol Agents
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2026)
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2026)
TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent
von: Meier, Dominik, et al.
Veröffentlicht: (2025)
von: Meier, Dominik, et al.
Veröffentlicht: (2025)
Block-wise Codeword Embedding for Reliable Multi-bit Text Watermarking
von: Kim, Joeun, et al.
Veröffentlicht: (2026)
von: Kim, Joeun, et al.
Veröffentlicht: (2026)
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents
von: Fu, Wenjie, et al.
Veröffentlicht: (2026)
von: Fu, Wenjie, et al.
Veröffentlicht: (2026)
StructuralSleight: Automated Jailbreak Attacks on Large Language Models Utilizing Uncommon Text-Organization Structures
von: Li, Bangxin, et al.
Veröffentlicht: (2024)
von: Li, Bangxin, et al.
Veröffentlicht: (2024)
Multi-Agent Collaboration in Incident Response with Large Language Models
von: Liu, Zefang
Veröffentlicht: (2024)
von: Liu, Zefang
Veröffentlicht: (2024)
Mitigating Cyber Risk in the Age of Open-Weight LLMs: Policy Gaps and Technical Realities
von: de Gregorio, Alfonso
Veröffentlicht: (2025)
von: de Gregorio, Alfonso
Veröffentlicht: (2025)
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
von: Zhang, Chiyu, et al.
Veröffentlicht: (2026)
von: Zhang, Chiyu, et al.
Veröffentlicht: (2026)
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
von: Wang, Peiran, et al.
Veröffentlicht: (2026)
von: Wang, Peiran, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs
von: Kim, Sangyeop, et al.
Veröffentlicht: (2025) -
Can LLM Infer Risk Information From MCP Server System Logs?
von: Fu, Jiayi, et al.
Veröffentlicht: (2025) -
Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback
von: Yan, Lecheng, et al.
Veröffentlicht: (2026) -
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
von: Lee, Hyomin, et al.
Veröffentlicht: (2026) -
LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops
von: Fu, Jiyuan, et al.
Veröffentlicht: (2025)