LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Zihe, Gui, Jiaping, Zhang, Zhuosheng, Liu, Gongshen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913008607821824
author Yan, Zihe
Gui, Jiaping
Zhang, Zhuosheng
Liu, Gongshen
author_facet Yan, Zihe
Gui, Jiaping
Zhang, Zhuosheng
Liu, Gongshen
contents Graphical user interface (GUI) agents built on multimodal large language models (MLLMs) have recently demonstrated strong decision-making abilities in screen-based interaction tasks. However, they remain highly vulnerable to pop-up-based environmental injection attacks, where malicious visual elements divert model attention and lead to unsafe or incorrect actions. Existing defense methods either require costly retraining or perform poorly under inductive interference. In this work, we systematically study how such attacks alter the attention behavior of GUI agents and uncover a layer-wise attention divergence pattern between correct and incorrect outputs. Based on this insight, we propose \textbf{LaSM}, a \textit{Layer-wise Scaling Mechanism} that selectively amplifies attention and MLP modules in critical layers. LaSM improves the alignment between model saliency and task-relevant regions without additional training. Extensive experiments across multiple datasets demonstrate that our method significantly improves the defense success rate and exhibits strong robustness, while having negligible impact on the model's general capabilities. Our findings reveal that attention misalignment is a core vulnerability in MLLM agents and can be effectively addressed through selective layer-wise modulation. Our code can be found in https://github.com/YANGTUOMAO/LaSM.
format Preprint
id arxiv_https___arxiv_org_abs_2507_10610
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI Agents
Yan, Zihe
Gui, Jiaping
Zhang, Zhuosheng
Liu, Gongshen
Cryptography and Security
Artificial Intelligence
Graphical user interface (GUI) agents built on multimodal large language models (MLLMs) have recently demonstrated strong decision-making abilities in screen-based interaction tasks. However, they remain highly vulnerable to pop-up-based environmental injection attacks, where malicious visual elements divert model attention and lead to unsafe or incorrect actions. Existing defense methods either require costly retraining or perform poorly under inductive interference. In this work, we systematically study how such attacks alter the attention behavior of GUI agents and uncover a layer-wise attention divergence pattern between correct and incorrect outputs. Based on this insight, we propose \textbf{LaSM}, a \textit{Layer-wise Scaling Mechanism} that selectively amplifies attention and MLP modules in critical layers. LaSM improves the alignment between model saliency and task-relevant regions without additional training. Extensive experiments across multiple datasets demonstrate that our method significantly improves the defense success rate and exhibits strong robustness, while having negligible impact on the model's general capabilities. Our findings reveal that attention misalignment is a core vulnerability in MLLM agents and can be effectively addressed through selective layer-wise modulation. Our code can be found in https://github.com/YANGTUOMAO/LaSM.
title LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI Agents
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2507.10610