Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Wenkui, Jin, Chao, Zhu, Haisu, Luo, Weilin, Yuen, Derek, Shao, Kun, Huang, Huaibo, Duan, Junxian, Cao, Jie, He, Ran |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Robust Defense against Customization via Protective Perturbation Resistant to Diffusion-based Purification
by: Yang, Wenkui, et al.
Published: (2025)
by: Yang, Wenkui, et al.
Published: (2025)
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
by: Qin, Yujia, et al.
Published: (2025)
by: Qin, Yujia, et al.
Published: (2025)
Visual Watermarking in the Era of Diffusion Models: Advances and Challenges
by: Duan, Junxian, et al.
Published: (2025)
by: Duan, Junxian, et al.
Published: (2025)
TT-DF: A Large-Scale Diffusion-Based Dataset and Benchmark for Human Body Forgery Detection
by: Yang, Wenkui, et al.
Published: (2025)
by: Yang, Wenkui, et al.
Published: (2025)
HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents
by: Jin, Chao, et al.
Published: (2026)
by: Jin, Chao, et al.
Published: (2026)
ShowUI-Aloha: Human-Taught GUI Agent
by: Zhang, Yichun, et al.
Published: (2026)
by: Zhang, Yichun, et al.
Published: (2026)
Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
by: Yang, Zhen, et al.
Published: (2025)
by: Yang, Zhen, et al.
Published: (2025)
Marmot: Object-Level Self-Correction via Multi-Agent Reasoning
by: Sun, Jiayang, et al.
Published: (2025)
by: Sun, Jiayang, et al.
Published: (2025)
Aria-UI: Visual Grounding for GUI Instructions
by: Yang, Yuhao, et al.
Published: (2024)
by: Yang, Yuhao, et al.
Published: (2024)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
Reducing Distraction in Long-Context Language Models by Focused Learning
by: Wu, Zijun, et al.
Published: (2024)
by: Wu, Zijun, et al.
Published: (2024)
ZePo: Zero-Shot Portrait Stylization with Faster Sampling
by: Liu, Jin, et al.
Published: (2024)
by: Liu, Jin, et al.
Published: (2024)
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
by: Wang, Haoming, et al.
Published: (2025)
by: Wang, Haoming, et al.
Published: (2025)
UFO: A UI-Focused Agent for Windows OS Interaction
by: Zhang, Chaoyun, et al.
Published: (2024)
by: Zhang, Chaoyun, et al.
Published: (2024)
BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
by: Zhang, Shaojie, et al.
Published: (2025)
by: Zhang, Shaojie, et al.
Published: (2025)
Automated UI Interface Generation via Diffusion Models: Enhancing Personalization and Efficiency
by: Duan, Yifei, et al.
Published: (2025)
by: Duan, Yifei, et al.
Published: (2025)
PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
by: Liu, Zikang, et al.
Published: (2025)
by: Liu, Zikang, et al.
Published: (2025)
MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
by: Zhou, Hanzhang, et al.
Published: (2025)
by: Zhou, Hanzhang, et al.
Published: (2025)
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
by: Sun, Qiushi, et al.
Published: (2024)
by: Sun, Qiushi, et al.
Published: (2024)
UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
by: Lian, Shuquan, et al.
Published: (2025)
by: Lian, Shuquan, et al.
Published: (2025)
Tuning Real-World Image Restoration at Inference: A Test-Time Scaling Paradigm for Flow Matching Models
by: Bai, Purui, et al.
Published: (2026)
by: Bai, Purui, et al.
Published: (2026)
Agent-Initiated Interaction in Phone UI Automation
by: Kahlon, Noam, et al.
Published: (2025)
by: Kahlon, Noam, et al.
Published: (2025)
AQuaUI: Visual Token Reduction for GUI Agents with Adaptive Quadtrees
by: Li, Yuankai, et al.
Published: (2026)
by: Li, Yuankai, et al.
Published: (2026)
UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience
by: Lin, Zichuan, et al.
Published: (2026)
by: Lin, Zichuan, et al.
Published: (2026)
UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding
by: Tang, Fei, et al.
Published: (2026)
by: Tang, Fei, et al.
Published: (2026)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
by: Jing, Hongyi, et al.
Published: (2025)
by: Jing, Hongyi, et al.
Published: (2025)
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
by: Xiao, Han, et al.
Published: (2025)
by: Xiao, Han, et al.
Published: (2025)
Falcon-UI: Understanding GUI Before Following User Instructions
by: Shen, Huawen, et al.
Published: (2024)
by: Shen, Huawen, et al.
Published: (2024)
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
by: Ouyang, Mingyu, et al.
Published: (2026)
by: Ouyang, Mingyu, et al.
Published: (2026)
Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens
by: Fan, Qihang, et al.
Published: (2024)
by: Fan, Qihang, et al.
Published: (2024)
NEXT: Multi-Grained Mixture of Experts via Text-Modulation for Multi-Modal Object Re-Identification
by: Li, Shihao, et al.
Published: (2025)
by: Li, Shihao, et al.
Published: (2025)
Breaking the Data Barrier -- Building GUI Agents Through Task Generalization
by: Zhang, Junlei, et al.
Published: (2025)
by: Zhang, Junlei, et al.
Published: (2025)
Vision Transformer with Super Token Sampling
by: Huang, Huaibo, et al.
Published: (2022)
by: Huang, Huaibo, et al.
Published: (2022)
TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents
by: Zhang, Bofei, et al.
Published: (2025)
by: Zhang, Bofei, et al.
Published: (2025)
UI-Mem: Self-Evolving Experience Memory for Online Reinforcement Learning in Mobile GUI Agents
by: Xiao, Han, et al.
Published: (2026)
by: Xiao, Han, et al.
Published: (2026)
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
by: Nayak, Shravan, et al.
Published: (2025)
by: Nayak, Shravan, et al.
Published: (2025)
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
by: Hu, Siyuan, et al.
Published: (2025)
by: Hu, Siyuan, et al.
Published: (2025)
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
by: Liu, Xinyi, et al.
Published: (2025)
by: Liu, Xinyi, et al.
Published: (2025)
The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections
by: Chen, Chaoran, et al.
Published: (2025)
by: Chen, Chaoran, et al.
Published: (2025)
Macaron-A2UI: A Model for Generative UI in Personal Agents
by: Kong, Fancy, et al.
Published: (2026)
by: Kong, Fancy, et al.
Published: (2026)
Similar Items
-
Towards Robust Defense against Customization via Protective Perturbation Resistant to Diffusion-based Purification
by: Yang, Wenkui, et al.
Published: (2025) -
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
by: Qin, Yujia, et al.
Published: (2025) -
Visual Watermarking in the Era of Diffusion Models: Advances and Challenges
by: Duan, Junxian, et al.
Published: (2025) -
TT-DF: A Large-Scale Diffusion-Based Dataset and Benchmark for Human Body Forgery Detection
by: Yang, Wenkui, et al.
Published: (2025) -
HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents
by: Jin, Chao, et al.
Published: (2026)