Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866918493067149312 |
|---|---|
| author | Zhang, Yilin Hua, Yingkai Wei, Chunyu Wang, Xin Chen, Yueguo |
| author_facet | Zhang, Yilin Hua, Yingkai Wei, Chunyu Wang, Xin Chen, Yueguo |
| contents | Vision-language model (VLM) based web agents demonstrate impressive autonomous GUI interaction but remain vulnerable to deceptive interface elements. Existing approaches either detect deception without task integration or document attacks without proposing defenses. We formalize deception-aware web agent defense and propose DUDE (Deceptive UI Detector & Evaluator), a two-stage framework combining hybrid-reward learning with asymmetric penalties and experience summarization to distill failure patterns into transferable guidance. We introduce RUC (Real UI Clickboxes), a benchmark of 1,407 scenarios spanning four domains and deception categories. Experiments show DUDE reduces deception susceptibility by 53.8% while maintaining task performance, establishing an effective foundation for robust web agent deployment. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_09497 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces Zhang, Yilin Hua, Yingkai Wei, Chunyu Wang, Xin Chen, Yueguo Artificial Intelligence Cryptography and Security Vision-language model (VLM) based web agents demonstrate impressive autonomous GUI interaction but remain vulnerable to deceptive interface elements. Existing approaches either detect deception without task integration or document attacks without proposing defenses. We formalize deception-aware web agent defense and propose DUDE (Deceptive UI Detector & Evaluator), a two-stage framework combining hybrid-reward learning with asymmetric penalties and experience summarization to distill failure patterns into transferable guidance. We introduce RUC (Real UI Clickboxes), a benchmark of 1,407 scenarios spanning four domains and deception categories. Experiments show DUDE reduces deception susceptibility by 53.8% while maintaining task performance, establishing an effective foundation for robust web agent deployment. |
| title | Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces |
| topic | Artificial Intelligence Cryptography and Security |
| url | https://arxiv.org/abs/2605.09497 |