Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | You, Keen, Zhang, Haotian, Schoop, Eldon, Weers, Floris, Swearngin, Amanda, Nichols, Jeffrey, Yang, Yinfei, Gan, Zhe |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Interaction to Impact: Towards Safer AI Agents Through Understanding and Evaluating Mobile UI Operation Impacts
di: Zhang, Zhuohao Jerry, et al.
Pubblicazione: (2024)
di: Zhang, Zhuohao Jerry, et al.
Pubblicazione: (2024)
Misty: UI Prototyping Through Interactive Conceptual Blending
di: Lu, Yuwen, et al.
Pubblicazione: (2024)
di: Lu, Yuwen, et al.
Pubblicazione: (2024)
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
di: Li, Zhangheng, et al.
Pubblicazione: (2024)
di: Li, Zhangheng, et al.
Pubblicazione: (2024)
AXNav: Replaying Accessibility Tests from Natural Language
di: Taeb, Maryam, et al.
Pubblicazione: (2023)
di: Taeb, Maryam, et al.
Pubblicazione: (2023)
Predicting and Explaining Mobile UI Tappability with Vision Modeling and Saliency Analysis
di: Schoop, Eldon, et al.
Pubblicazione: (2022)
di: Schoop, Eldon, et al.
Pubblicazione: (2022)
Generative UI: LLMs are Effective UI Generators
di: Leviathan, Yaniv, et al.
Pubblicazione: (2026)
di: Leviathan, Yaniv, et al.
Pubblicazione: (2026)
Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations
di: Liang, Chen, et al.
Pubblicazione: (2026)
di: Liang, Chen, et al.
Pubblicazione: (2026)
UICoder: Finetuning Large Language Models to Generate User Interface Code through Automated Feedback
di: Wu, Jason, et al.
Pubblicazione: (2024)
di: Wu, Jason, et al.
Pubblicazione: (2024)
Mapping the Design Space of User Experience for Computer Use Agents
di: Cheng, Ruijia, et al.
Pubblicazione: (2026)
di: Cheng, Ruijia, et al.
Pubblicazione: (2026)
Athena: Intermediate Representations for Iterative Scaffolded App Generation with an LLM
di: Beason, Jazbo, et al.
Pubblicazione: (2025)
di: Beason, Jazbo, et al.
Pubblicazione: (2025)
MAIC-UI: Making Interactive Courseware with Generative UI
di: Tu, Shangqing, et al.
Pubblicazione: (2026)
di: Tu, Shangqing, et al.
Pubblicazione: (2026)
ReDemon UI: Reactive Synthesis by Demonstration for Web UI
di: Lee, Jay, et al.
Pubblicazione: (2025)
di: Lee, Jay, et al.
Pubblicazione: (2025)
GhostUI: Unveiling Hidden Interactions in Mobile UI
di: Kweon, Minkyu, et al.
Pubblicazione: (2026)
di: Kweon, Minkyu, et al.
Pubblicazione: (2026)
Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
di: Yang, Zhen, et al.
Pubblicazione: (2025)
di: Yang, Zhen, et al.
Pubblicazione: (2025)
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
di: Ouyang, Mingyu, et al.
Pubblicazione: (2026)
di: Ouyang, Mingyu, et al.
Pubblicazione: (2026)
UIClip: A Data-driven Model for Assessing User Interface Design
di: Wu, Jason, et al.
Pubblicazione: (2024)
di: Wu, Jason, et al.
Pubblicazione: (2024)
UI-Evol: Automatic Knowledge Evolving for Computer Use Agents
di: Zhang, Ziyun, et al.
Pubblicazione: (2025)
di: Zhang, Ziyun, et al.
Pubblicazione: (2025)
Morae: Proactively Pausing UI Agents for User Choices
di: Peng, Yi-Hao, et al.
Pubblicazione: (2025)
di: Peng, Yi-Hao, et al.
Pubblicazione: (2025)
The Way We Notice, That's What Really Matters: Instantiating UI Components with Distinguishing Variations
di: Vaithilingam, Priyan, et al.
Pubblicazione: (2026)
di: Vaithilingam, Priyan, et al.
Pubblicazione: (2026)
RWKV-UI: UI Understanding with Enhanced Perception and Reasoning
di: Yang, Jiaxi, et al.
Pubblicazione: (2025)
di: Yang, Jiaxi, et al.
Pubblicazione: (2025)
MUIAnno: An Expert-Annotated Dataset and Evaluation Benchmark for Mobile UI Understanding
di: Parvez, Athar, et al.
Pubblicazione: (2026)
di: Parvez, Athar, et al.
Pubblicazione: (2026)
User-Centric Design of UI for Mobile Banking Apps: Improving UI and Features for Better Customer Experience
di: Chitrakar, Luniva, et al.
Pubblicazione: (2026)
di: Chitrakar, Luniva, et al.
Pubblicazione: (2026)
FlowEval: Reference-based Evaluation of Generated User Interfaces
di: Wu, Jason, et al.
Pubblicazione: (2026)
di: Wu, Jason, et al.
Pubblicazione: (2026)
AgentBuilder: Exploring Scaffolds for Prototyping User Experiences of Interface Agents
di: Liang, Jenny T., et al.
Pubblicazione: (2025)
di: Liang, Jenny T., et al.
Pubblicazione: (2025)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
di: Jing, Hongyi, et al.
Pubblicazione: (2025)
di: Jing, Hongyi, et al.
Pubblicazione: (2025)
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
di: Liu, Xinyi, et al.
Pubblicazione: (2025)
di: Liu, Xinyi, et al.
Pubblicazione: (2025)
UFO: A UI-Focused Agent for Windows OS Interaction
di: Zhang, Chaoyun, et al.
Pubblicazione: (2024)
di: Zhang, Chaoyun, et al.
Pubblicazione: (2024)
UI-UG: A Unified MLLM for UI Understanding and Generation
di: Yang, Hao, et al.
Pubblicazione: (2025)
di: Yang, Hao, et al.
Pubblicazione: (2025)
AURORA: Navigating UI Tarpits via Automated Neural Screen Understanding
di: Khan, Safwat Ali, et al.
Pubblicazione: (2024)
di: Khan, Safwat Ali, et al.
Pubblicazione: (2024)
Privacy Starts with UI: Privacy Patterns and Designer Perspectives in UI/UX Practice
di: Maloku, Anxhela, et al.
Pubblicazione: (2026)
di: Maloku, Anxhela, et al.
Pubblicazione: (2026)
Macaron-A2UI: A Model for Generative UI in Personal Agents
di: Kong, Fancy, et al.
Pubblicazione: (2026)
di: Kong, Fancy, et al.
Pubblicazione: (2026)
UI Remix: Supporting UI Design Through Interactive Example Retrieval and Remixing
di: Wang, Junling, et al.
Pubblicazione: (2026)
di: Wang, Junling, et al.
Pubblicazione: (2026)
CrowdGenUI: Aligning LLM-Based UI Generation with Crowdsourced User Preferences
di: Liu, Yimeng, et al.
Pubblicazione: (2024)
di: Liu, Yimeng, et al.
Pubblicazione: (2024)
VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task Planning
di: Song, Yunpeng, et al.
Pubblicazione: (2023)
di: Song, Yunpeng, et al.
Pubblicazione: (2023)
Context-Aware Workflow Decomposition for Automated Mobile UI Annotation Using Multimodal Large Language Models
di: Parvez, Athar, et al.
Pubblicazione: (2026)
di: Parvez, Athar, et al.
Pubblicazione: (2026)
Affordances of Sketched Notations for Multimodal UI Design and Development Tools
di: Ross, Sam H., et al.
Pubblicazione: (2025)
di: Ross, Sam H., et al.
Pubblicazione: (2025)
UISim: An Interactive Image-Based UI Simulator for Dynamic Mobile Environments
di: Xiang, Jiannan, et al.
Pubblicazione: (2025)
di: Xiang, Jiannan, et al.
Pubblicazione: (2025)
Aria-UI: Visual Grounding for GUI Instructions
di: Yang, Yuhao, et al.
Pubblicazione: (2024)
di: Yang, Yuhao, et al.
Pubblicazione: (2024)
AutoGameUI: Constructing High-Fidelity GameUI via Multimodal Correspondence Matching
di: Tang, Zhongliang, et al.
Pubblicazione: (2024)
di: Tang, Zhongliang, et al.
Pubblicazione: (2024)
The GenUI Study: Exploring the Design of Generative UI Tools to Support UX Practitioners and Beyond
di: Chen, Xiang 'Anthony', et al.
Pubblicazione: (2025)
di: Chen, Xiang 'Anthony', et al.
Pubblicazione: (2025)
Documenti analoghi
-
From Interaction to Impact: Towards Safer AI Agents Through Understanding and Evaluating Mobile UI Operation Impacts
di: Zhang, Zhuohao Jerry, et al.
Pubblicazione: (2024) -
Misty: UI Prototyping Through Interactive Conceptual Blending
di: Lu, Yuwen, et al.
Pubblicazione: (2024) -
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
di: Li, Zhangheng, et al.
Pubblicazione: (2024) -
AXNav: Replaying Accessibility Tests from Natural Language
di: Taeb, Maryam, et al.
Pubblicazione: (2023) -
Predicting and Explaining Mobile UI Tappability with Vision Modeling and Saliency Analysis
di: Schoop, Eldon, et al.
Pubblicazione: (2022)