Hallucination as Exploit: Evidence-Carrying Multimodal Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Guijia, Zheng, Hao, Yang, Harry |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI Agent Smart Contract Exploit Generation
by: Gervais, Arthur, et al.
Published: (2025)
by: Gervais, Arthur, et al.
Published: (2025)
When Child Inherits: Modeling and Exploiting Subagent Spawn in Multi-Agent Networks
by: Cai, Ziwen, et al.
Published: (2026)
by: Cai, Ziwen, et al.
Published: (2026)
Atomicity for Agents: Exposing, Exploiting, and Mitigating TOCTOU Vulnerabilities in Browser-Use Agents
by: Jiang, Linxi, et al.
Published: (2026)
by: Jiang, Linxi, et al.
Published: (2026)
LLM Agents can Autonomously Exploit One-day Vulnerabilities
by: Fang, Richard, et al.
Published: (2024)
by: Fang, Richard, et al.
Published: (2024)
CompressionAttack: Exploiting Prompt Compression as a New Attack Surface in LLM-Powered Agents
by: Liu, Zesen, et al.
Published: (2025)
by: Liu, Zesen, et al.
Published: (2025)
ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
by: Lee, Seunghyun, et al.
Published: (2026)
by: Lee, Seunghyun, et al.
Published: (2026)
Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents
by: Zou, Wei, et al.
Published: (2026)
by: Zou, Wei, et al.
Published: (2026)
From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
CybORG++: An Enhanced Gym for the Development of Autonomous Cyber Agents
by: Emerson, Harry, et al.
Published: (2024)
by: Emerson, Harry, et al.
Published: (2024)
DECEPTICON: How Dark Patterns Manipulate Web Agents
by: Cuvin, Phil, et al.
Published: (2025)
by: Cuvin, Phil, et al.
Published: (2025)
AgentSys: Secure and Dynamic LLM Agents Through Explicit Hierarchical Memory Management
by: Wen, Ruoyao, et al.
Published: (2026)
by: Wen, Ruoyao, et al.
Published: (2026)
Exploiting Prefix-Tree in Structured Output Interfaces for Enhancing Jailbreak Attacking
by: Li, Yanzeng, et al.
Published: (2025)
by: Li, Yanzeng, et al.
Published: (2025)
MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation
by: Yan, Lu, et al.
Published: (2025)
by: Yan, Lu, et al.
Published: (2025)
AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
by: Li, Yanjie, et al.
Published: (2025)
by: Li, Yanjie, et al.
Published: (2025)
LLM Safeguard is a Double-Edged Sword: Exploiting False Positives for Denial-of-Service Attacks
by: Zhang, Qingzhao, et al.
Published: (2024)
by: Zhang, Qingzhao, et al.
Published: (2024)
AudAgent: Automated Auditing of Privacy Policy Compliance in AI Agents
by: Zheng, Ye, et al.
Published: (2025)
by: Zheng, Ye, et al.
Published: (2025)
CORVUS: Red-Teaming Hallucination Detectors via Internal Signal Camouflage in Large Language Models
by: Min, Nay Myat, et al.
Published: (2026)
by: Min, Nay Myat, et al.
Published: (2026)
Prompt-in-Content Attacks: Exploiting Uploaded Inputs to Hijack LLM Behavior
by: Lian, Zhuotao, et al.
Published: (2025)
by: Lian, Zhuotao, et al.
Published: (2025)
Hallucination-Resistant Security Planning with a Large Language Model
by: Hammar, Kim, et al.
Published: (2026)
by: Hammar, Kim, et al.
Published: (2026)
MTRE: Multi-Token Reliability Estimation for Hallucination Detection in VLMs
by: Zollicoffer, Geigh, et al.
Published: (2025)
by: Zollicoffer, Geigh, et al.
Published: (2025)
SafeMobile: Chain-level Jailbreak Detection and Automated Evaluation for Multimodal Mobile Agents
by: Liang, Siyuan, et al.
Published: (2025)
by: Liang, Siyuan, et al.
Published: (2025)
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
by: Wang, Zhun, et al.
Published: (2026)
by: Wang, Zhun, et al.
Published: (2026)
LegalSim: Multi-Agent Simulation of Legal Systems for Discovering Procedural Exploits
by: Badhe, Sanket
Published: (2025)
by: Badhe, Sanket
Published: (2025)
Security of AI Agents
by: He, Yifeng, et al.
Published: (2024)
by: He, Yifeng, et al.
Published: (2024)
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
by: Zheng, Baolin, et al.
Published: (2025)
by: Zheng, Baolin, et al.
Published: (2025)
AgentMark: Utility-Preserving Behavioral Watermarking for Agents
by: Huang, Kaibo, et al.
Published: (2026)
by: Huang, Kaibo, et al.
Published: (2026)
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use
by: Yang, Chenglin
Published: (2026)
by: Yang, Chenglin
Published: (2026)
Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations
by: Collu, Matteo Gioele, et al.
Published: (2026)
by: Collu, Matteo Gioele, et al.
Published: (2026)
"Glue pizza and eat rocks" -- Exploiting Vulnerabilities in Retrieval-Augmented Generative Models
by: Tan, Zhen, et al.
Published: (2024)
by: Tan, Zhen, et al.
Published: (2024)
Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations
by: Wong, Ryan, et al.
Published: (2025)
by: Wong, Ryan, et al.
Published: (2025)
Audit-LLM: Multi-Agent Collaboration for Log-based Insider Threat Detection
by: Song, Chengyu, et al.
Published: (2024)
by: Song, Chengyu, et al.
Published: (2024)
MANA: Towards Efficient Mobile Ad Detection via Multimodal Agentic UI Navigation
by: Zhao, Yizhe, et al.
Published: (2026)
by: Zhao, Yizhe, et al.
Published: (2026)
Good News for Script Kiddies? Evaluating Large Language Models for Automated Exploit Generation
by: Jin, David, et al.
Published: (2025)
by: Jin, David, et al.
Published: (2025)
Hallucinating AI Hijacking Attack: Large Language Models and Malicious Code Recommenders
by: Noever, David, et al.
Published: (2024)
by: Noever, David, et al.
Published: (2024)
Exploiting Efficiency Vulnerabilities in Dynamic Deep Learning Systems
by: Rathnasuriya, Ravishka, et al.
Published: (2025)
by: Rathnasuriya, Ravishka, et al.
Published: (2025)
DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs
by: Xu, Wenzhuo, et al.
Published: (2026)
by: Xu, Wenzhuo, et al.
Published: (2026)
ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety
by: Wang, Kun, et al.
Published: (2026)
by: Wang, Kun, et al.
Published: (2026)
CheatAgent: Attacking LLM-Empowered Recommender Systems via LLM Agent
by: Ning, Liang-bo, et al.
Published: (2025)
by: Ning, Liang-bo, et al.
Published: (2025)
Compiled Models, Built-In Exploits: Uncovering Pervasive Bit-Flip Attack Surfaces in DNN Executables
by: Chen, Yanzuo, et al.
Published: (2023)
by: Chen, Yanzuo, et al.
Published: (2023)
MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMs
by: Ding, Ruyi, et al.
Published: (2025)
by: Ding, Ruyi, et al.
Published: (2025)
Similar Items
-
AI Agent Smart Contract Exploit Generation
by: Gervais, Arthur, et al.
Published: (2025) -
When Child Inherits: Modeling and Exploiting Subagent Spawn in Multi-Agent Networks
by: Cai, Ziwen, et al.
Published: (2026) -
Atomicity for Agents: Exposing, Exploiting, and Mitigating TOCTOU Vulnerabilities in Browser-Use Agents
by: Jiang, Linxi, et al.
Published: (2026) -
LLM Agents can Autonomously Exploit One-day Vulnerabilities
by: Fang, Richard, et al.
Published: (2024) -
CompressionAttack: Exploiting Prompt Compression as a New Attack Surface in LLM-Powered Agents
by: Liu, Zesen, et al.
Published: (2025)