Zero-Permission Manipulation: Can We Trust Large Multimodal Model Powered GUI Agents?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Qian, Yi, Qian, Kunwei, He, Xingbang, Chen, Ligeng, Zhang, Jikang, Zhang, Tiantai, Wei, Haiyang, Wang, Linzhang, Wu, Hao, Mao, Bing
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915829661040640
author Qian, Yi
Qian, Kunwei
He, Xingbang
Chen, Ligeng
Zhang, Jikang
Zhang, Tiantai
Wei, Haiyang
Wang, Linzhang
Wu, Hao
Mao, Bing
author_facet Qian, Yi
Qian, Kunwei
He, Xingbang
Chen, Ligeng
Zhang, Jikang
Zhang, Tiantai
Wei, Haiyang
Wang, Linzhang
Wu, Hao
Mao, Bing
contents Large multimodal model powered GUI agents are emerging as high-privilege operators on mobile platforms, entrusted with perceiving screen content and injecting inputs. However, their design operates under the implicit assumption of Visual Atomicity: that the UI state remains invariant between observation and action. We demonstrate that this assumption is fundamentally invalid in Android, creating a critical attack surface. We present Action Rebinding, a novel attack that allows a seemingly-benign app with zero dangerous permissions to rebind an agent's execution. By exploiting the inevitable observation-to-action gap inherent in the agent's reasoning pipeline, the attacker triggers foreground transitions to rebind the agent's planned action toward the target app. We weaponize the agent's task-recovery logic and Android's UI state preservation to orchestrate programmable, multi-step attack chains. Furthermore, we introduce an Intent Alignment Strategy (IAS) that manipulates the agent's reasoning process to rationalize UI states, enabling it to bypass verification gates (e.g., confirmation dialogs) that would otherwise be rejected. We evaluate Action Rebinding Attacks on six widely-used Android GUI agents across 15 tasks. Our results demonstrate a 100% success rate for atomic action rebinding and the ability to reliably orchestrate multi-step attack chains. With IAS, the success rate in bypassing verification gates increases (from 0% to up to 100%). Notably, the attacker application requires no sensitive permissions and contains no privileged API calls, achieving a 0% detection rate across malware scanners (e.g., VirusTotal). Our findings reveal a fundamental architectural flaw in current agent-OS integration and provide critical insights for the secure design of future agent systems. To access experimental logs and demonstration videos, please contact yi_qian@smail.nju.edu.cn.
format Preprint
id arxiv_https___arxiv_org_abs_2601_12349
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Zero-Permission Manipulation: Can We Trust Large Multimodal Model Powered GUI Agents?
Qian, Yi
Qian, Kunwei
He, Xingbang
Chen, Ligeng
Zhang, Jikang
Zhang, Tiantai
Wei, Haiyang
Wang, Linzhang
Wu, Hao
Mao, Bing
Cryptography and Security
Artificial Intelligence
Software Engineering
Large multimodal model powered GUI agents are emerging as high-privilege operators on mobile platforms, entrusted with perceiving screen content and injecting inputs. However, their design operates under the implicit assumption of Visual Atomicity: that the UI state remains invariant between observation and action. We demonstrate that this assumption is fundamentally invalid in Android, creating a critical attack surface. We present Action Rebinding, a novel attack that allows a seemingly-benign app with zero dangerous permissions to rebind an agent's execution. By exploiting the inevitable observation-to-action gap inherent in the agent's reasoning pipeline, the attacker triggers foreground transitions to rebind the agent's planned action toward the target app. We weaponize the agent's task-recovery logic and Android's UI state preservation to orchestrate programmable, multi-step attack chains. Furthermore, we introduce an Intent Alignment Strategy (IAS) that manipulates the agent's reasoning process to rationalize UI states, enabling it to bypass verification gates (e.g., confirmation dialogs) that would otherwise be rejected. We evaluate Action Rebinding Attacks on six widely-used Android GUI agents across 15 tasks. Our results demonstrate a 100% success rate for atomic action rebinding and the ability to reliably orchestrate multi-step attack chains. With IAS, the success rate in bypassing verification gates increases (from 0% to up to 100%). Notably, the attacker application requires no sensitive permissions and contains no privileged API calls, achieving a 0% detection rate across malware scanners (e.g., VirusTotal). Our findings reveal a fundamental architectural flaw in current agent-OS integration and provide critical insights for the secure design of future agent systems. To access experimental logs and demonstration videos, please contact yi_qian@smail.nju.edu.cn.
title Zero-Permission Manipulation: Can We Trust Large Multimodal Model Powered GUI Agents?
topic Cryptography and Security
Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2601.12349