FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
Fuente:
arXiv
Saved in:
| Main Authors: | Ouyang, Mingyu, Lin, Kevin Qinghong, Shou, Mike Zheng, Ng, Hwee Tou |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
by: Ouyang, Mingyu, et al.
Published: (2026)
by: Ouyang, Mingyu, et al.
Published: (2026)
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
by: Hu, Siyuan, et al.
Published: (2025)
by: Hu, Siyuan, et al.
Published: (2025)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
by: You, Keen, et al.
Published: (2024)
by: You, Keen, et al.
Published: (2024)
Aria-UI: Visual Grounding for GUI Instructions
by: Yang, Yuhao, et al.
Published: (2024)
by: Yang, Yuhao, et al.
Published: (2024)
Macaron-A2UI: A Model for Generative UI in Personal Agents
by: Kong, Fancy, et al.
Published: (2026)
by: Kong, Fancy, et al.
Published: (2026)
Privacy Starts with UI: Privacy Patterns and Designer Perspectives in UI/UX Practice
by: Maloku, Anxhela, et al.
Published: (2026)
by: Maloku, Anxhela, et al.
Published: (2026)
UI Remix: Supporting UI Design Through Interactive Example Retrieval and Remixing
by: Wang, Junling, et al.
Published: (2026)
by: Wang, Junling, et al.
Published: (2026)
UFO: A UI-Focused Agent for Windows OS Interaction
by: Zhang, Chaoyun, et al.
Published: (2024)
by: Zhang, Chaoyun, et al.
Published: (2024)
ProactiveVA: Proactive Visual Analytics with LLM-Based UI Agent
by: Zhao, Yuheng, et al.
Published: (2025)
by: Zhao, Yuheng, et al.
Published: (2025)
CrowdGenUI: Aligning LLM-Based UI Generation with Crowdsourced User Preferences
by: Liu, Yimeng, et al.
Published: (2024)
by: Liu, Yimeng, et al.
Published: (2024)
RWKV-UI: UI Understanding with Enhanced Perception and Reasoning
by: Yang, Jiaxi, et al.
Published: (2025)
by: Yang, Jiaxi, et al.
Published: (2025)
MAIC-UI: Making Interactive Courseware with Generative UI
by: Tu, Shangqing, et al.
Published: (2026)
by: Tu, Shangqing, et al.
Published: (2026)
The GenUI Study: Exploring the Design of Generative UI Tools to Support UX Practitioners and Beyond
by: Chen, Xiang 'Anthony', et al.
Published: (2025)
by: Chen, Xiang 'Anthony', et al.
Published: (2025)
GhostUI: Unveiling Hidden Interactions in Mobile UI
by: Kweon, Minkyu, et al.
Published: (2026)
by: Kweon, Minkyu, et al.
Published: (2026)
DynaVis: Dynamically Synthesized UI Widgets for Visualization Editing
by: Vaithilingam, Priyan, et al.
Published: (2024)
by: Vaithilingam, Priyan, et al.
Published: (2024)
SpecifyUI: Supporting Iterative UI Design Intent Expression through Structured Specifications and Generative AI
by: Chen, Yunnong, et al.
Published: (2025)
by: Chen, Yunnong, et al.
Published: (2025)
User-Centric Design of UI for Mobile Banking Apps: Improving UI and Features for Better Customer Experience
by: Chitrakar, Luniva, et al.
Published: (2026)
by: Chitrakar, Luniva, et al.
Published: (2026)
MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI Modeling
by: Feng, Sidong, et al.
Published: (2024)
by: Feng, Sidong, et al.
Published: (2024)
What is the focus of XAI in UI design? Prioritizing UI design principles for enhancing XAI user experience
by: Lei, Dian, et al.
Published: (2024)
by: Lei, Dian, et al.
Published: (2024)
ReDemon UI: Reactive Synthesis by Demonstration for Web UI
by: Lee, Jay, et al.
Published: (2025)
by: Lee, Jay, et al.
Published: (2025)
Generative UI: LLMs are Effective UI Generators
by: Leviathan, Yaniv, et al.
Published: (2026)
by: Leviathan, Yaniv, et al.
Published: (2026)
UI-UG: A Unified MLLM for UI Understanding and Generation
by: Yang, Hao, et al.
Published: (2025)
by: Yang, Hao, et al.
Published: (2025)
AutoGameUI: Constructing High-Fidelity GameUI via Multimodal Correspondence Matching
by: Tang, Zhongliang, et al.
Published: (2024)
by: Tang, Zhongliang, et al.
Published: (2024)
Code2Video: A Code-centric Paradigm for Educational Video Generation
by: Chen, Yanzhe, et al.
Published: (2025)
by: Chen, Yanzhe, et al.
Published: (2025)
Preference-Guided Multi-Objective UI Adaptation
by: Song, Yao, et al.
Published: (2025)
by: Song, Yao, et al.
Published: (2025)
Agent-Initiated Interaction in Phone UI Automation
by: Kahlon, Noam, et al.
Published: (2025)
by: Kahlon, Noam, et al.
Published: (2025)
MLLM-Based UI2Code Automation Guided by UI Layout Information
by: Wu, Fan, et al.
Published: (2025)
by: Wu, Fan, et al.
Published: (2025)
SuperProvenanceWidgets: Tracking and Visualizing Analytic Provenance Across UI Control Elements
by: Verma, Antariksh, et al.
Published: (2026)
by: Verma, Antariksh, et al.
Published: (2026)
Misty: UI Prototyping Through Interactive Conceptual Blending
by: Lu, Yuwen, et al.
Published: (2024)
by: Lu, Yuwen, et al.
Published: (2024)
Automating UI Optimization through Multi-Agentic Reasoning
by: Li, Zhipeng, et al.
Published: (2026)
by: Li, Zhipeng, et al.
Published: (2026)
Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations
by: Liang, Chen, et al.
Published: (2026)
by: Liang, Chen, et al.
Published: (2026)
Automated UI Interface Generation via Diffusion Models: Enhancing Personalization and Efficiency
by: Duan, Yifei, et al.
Published: (2025)
by: Duan, Yifei, et al.
Published: (2025)
Generating Automatic Feedback on UI Mockups with Large Language Models
by: Duan, Peitong, et al.
Published: (2024)
by: Duan, Peitong, et al.
Published: (2024)
Affordances of Sketched Notations for Multimodal UI Design and Development Tools
by: Ross, Sam H., et al.
Published: (2025)
by: Ross, Sam H., et al.
Published: (2025)
UI-Evol: Automatic Knowledge Evolving for Computer Use Agents
by: Zhang, Ziyun, et al.
Published: (2025)
by: Zhang, Ziyun, et al.
Published: (2025)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
by: Jing, Hongyi, et al.
Published: (2025)
by: Jing, Hongyi, et al.
Published: (2025)
Efficient and Aesthetic UI Design with a Deep Learning-Based Interface Generation Tree Algorithm
by: Duan, Shiyu, et al.
Published: (2024)
by: Duan, Shiyu, et al.
Published: (2024)
MUIAnno: An Expert-Annotated Dataset and Evaluation Benchmark for Mobile UI Understanding
by: Parvez, Athar, et al.
Published: (2026)
by: Parvez, Athar, et al.
Published: (2026)
Open WebUI: An Open, Extensible, and Usable Interface for AI Interaction
by: Baek, Jaeryang, et al.
Published: (2025)
by: Baek, Jaeryang, et al.
Published: (2025)
Similar Items
-
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
by: Ouyang, Mingyu, et al.
Published: (2026) -
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
by: Hu, Siyuan, et al.
Published: (2025) -
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
by: Lin, Kevin Qinghong, et al.
Published: (2024) -
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
by: You, Keen, et al.
Published: (2024) -
Aria-UI: Visual Grounding for GUI Instructions
by: Yang, Yuhao, et al.
Published: (2024)