Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Yilin, Hua, Yingkai, Wei, Chunyu, Wang, Xin, Chen, Yueguo
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918493067149312
author Zhang, Yilin
Hua, Yingkai
Wei, Chunyu
Wang, Xin
Chen, Yueguo
author_facet Zhang, Yilin
Hua, Yingkai
Wei, Chunyu
Wang, Xin
Chen, Yueguo
contents Vision-language model (VLM) based web agents demonstrate impressive autonomous GUI interaction but remain vulnerable to deceptive interface elements. Existing approaches either detect deception without task integration or document attacks without proposing defenses. We formalize deception-aware web agent defense and propose DUDE (Deceptive UI Detector & Evaluator), a two-stage framework combining hybrid-reward learning with asymmetric penalties and experience summarization to distill failure patterns into transferable guidance. We introduce RUC (Real UI Clickboxes), a benchmark of 1,407 scenarios spanning four domains and deception categories. Experiments show DUDE reduces deception susceptibility by 53.8% while maintaining task performance, establishing an effective foundation for robust web agent deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2605_09497
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces
Zhang, Yilin
Hua, Yingkai
Wei, Chunyu
Wang, Xin
Chen, Yueguo
Artificial Intelligence
Cryptography and Security
Vision-language model (VLM) based web agents demonstrate impressive autonomous GUI interaction but remain vulnerable to deceptive interface elements. Existing approaches either detect deception without task integration or document attacks without proposing defenses. We formalize deception-aware web agent defense and propose DUDE (Deceptive UI Detector & Evaluator), a two-stage framework combining hybrid-reward learning with asymmetric penalties and experience summarization to distill failure patterns into transferable guidance. We introduce RUC (Real UI Clickboxes), a benchmark of 1,407 scenarios spanning four domains and deception categories. Experiments show DUDE reduces deception susceptibility by 53.8% while maintaining task performance, establishing an effective foundation for robust web agent deployment.
title Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces
topic Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2605.09497