Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zhangheng, You, Keen, Zhang, Haotian, Feng, Di, Agrawal, Harsh, Li, Xiujun, Moorthy, Mohana Prasad Sathya, Nichols, Jeff, Yang, Yinfei, Gan, Zhe |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
by: You, Keen, et al.
Published: (2024)
by: You, Keen, et al.
Published: (2024)
Prioritising Interactive Flows in Data Center Networks With Central Control
by: Moorthy, Mohana Prasad Sathya
Published: (2023)
by: Moorthy, Mohana Prasad Sathya
Published: (2023)
Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
by: Yang, Zhen, et al.
Published: (2025)
by: Yang, Zhen, et al.
Published: (2025)
Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
by: Zhang, Haotian, et al.
Published: (2024)
by: Zhang, Haotian, et al.
Published: (2024)
How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
by: Qian, Yusu, et al.
Published: (2024)
by: Qian, Yusu, et al.
Published: (2024)
Contrastive Localized Language-Image Pre-Training
by: Chen, Hong-You, et al.
Published: (2024)
by: Chen, Hong-You, et al.
Published: (2024)
Morae: Proactively Pausing UI Agents for User Choices
by: Peng, Yi-Hao, et al.
Published: (2025)
by: Peng, Yi-Hao, et al.
Published: (2025)
UIClip: A Data-driven Model for Assessing User Interface Design
by: Wu, Jason, et al.
Published: (2024)
by: Wu, Jason, et al.
Published: (2024)
Hidden Technical Debt in Generative (GenUI) and Malleable User Interfaces
by: Cifliku, Besjon
Published: (2026)
by: Cifliku, Besjon
Published: (2026)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
by: Zang, Yuan, et al.
Published: (2025)
by: Zang, Yuan, et al.
Published: (2025)
Text as Images: Can Multimodal Large Language Models Follow Printed Instructions in Pixels?
by: Li, Xiujun, et al.
Published: (2023)
by: Li, Xiujun, et al.
Published: (2023)
UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
by: Yang, Yuhao, et al.
Published: (2025)
by: Yang, Yuhao, et al.
Published: (2025)
MAIC-UI: Making Interactive Courseware with Generative UI
by: Tu, Shangqing, et al.
Published: (2026)
by: Tu, Shangqing, et al.
Published: (2026)
Misty: UI Prototyping Through Interactive Conceptual Blending
by: Lu, Yuwen, et al.
Published: (2024)
by: Lu, Yuwen, et al.
Published: (2024)
Falcon-UI: Understanding GUI Before Following User Instructions
by: Shen, Huawen, et al.
Published: (2024)
by: Shen, Huawen, et al.
Published: (2024)
Natural Language Interfaces for Databases: What Do Users Think?
by: Ipeirotis, Panos, et al.
Published: (2025)
by: Ipeirotis, Panos, et al.
Published: (2025)
MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
by: Qian, Yusu, et al.
Published: (2024)
by: Qian, Yusu, et al.
Published: (2024)
CrowdGenUI: Aligning LLM-Based UI Generation with Crowdsourced User Preferences
by: Liu, Yimeng, et al.
Published: (2024)
by: Liu, Yimeng, et al.
Published: (2024)
MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces
by: Luera, Reuben A., et al.
Published: (2025)
by: Luera, Reuben A., et al.
Published: (2025)
Improve Vision Language Model Chain-of-thought Reasoning
by: Zhang, Ruohong, et al.
Published: (2024)
by: Zhang, Ruohong, et al.
Published: (2024)
The Way We Notice, That's What Really Matters: Instantiating UI Components with Distinguishing Variations
by: Vaithilingam, Priyan, et al.
Published: (2026)
by: Vaithilingam, Priyan, et al.
Published: (2026)
Hector UI: A Flexible Human-Robot User Interface for (Semi-)Autonomous Rescue and Inspection Robots
by: Fabian, Stefan, et al.
Published: (2025)
by: Fabian, Stefan, et al.
Published: (2025)
GhostUI: Unveiling Hidden Interactions in Mobile UI
by: Kweon, Minkyu, et al.
Published: (2026)
by: Kweon, Minkyu, et al.
Published: (2026)
Identifying User Goals from UI Trajectories
by: Berkovitch, Omri, et al.
Published: (2024)
by: Berkovitch, Omri, et al.
Published: (2024)
User-Centric Design of UI for Mobile Banking Apps: Improving UI and Features for Better Customer Experience
by: Chitrakar, Luniva, et al.
Published: (2026)
by: Chitrakar, Luniva, et al.
Published: (2026)
PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
by: Qian, Yusu, et al.
Published: (2025)
by: Qian, Yusu, et al.
Published: (2025)
Compressing LLMs: The Truth is Rarely Pure and Never Simple
by: Jaiswal, Ajay, et al.
Published: (2023)
by: Jaiswal, Ajay, et al.
Published: (2023)
FlowEval: Reference-based Evaluation of Generated User Interfaces
by: Wu, Jason, et al.
Published: (2026)
by: Wu, Jason, et al.
Published: (2026)
MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
by: Zhang, Haotian, et al.
Published: (2024)
by: Zhang, Haotian, et al.
Published: (2024)
UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity
by: Fu, Yicheng, et al.
Published: (2024)
by: Fu, Yicheng, et al.
Published: (2024)
MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA
by: Ye, Hanrong, et al.
Published: (2024)
by: Ye, Hanrong, et al.
Published: (2024)
AgentBuilder: Exploring Scaffolds for Prototyping User Experiences of Interface Agents
by: Liang, Jenny T., et al.
Published: (2025)
by: Liang, Jenny T., et al.
Published: (2025)
UICoder: Finetuning Large Language Models to Generate User Interface Code through Automated Feedback
by: Wu, Jason, et al.
Published: (2024)
by: Wu, Jason, et al.
Published: (2024)
DP-OPT: Make Large Language Model Your Privacy-Preserving Prompt Engineer
by: Hong, Junyuan, et al.
Published: (2023)
by: Hong, Junyuan, et al.
Published: (2023)
Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design Understanding
by: Jeon, Jaehyun, et al.
Published: (2025)
by: Jeon, Jaehyun, et al.
Published: (2025)
UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing
by: Fu, Tsu-Jui, et al.
Published: (2025)
by: Fu, Tsu-Jui, et al.
Published: (2025)
Improving User Interface Generation Models from Designer Feedback
by: Wu, Jason, et al.
Published: (2025)
by: Wu, Jason, et al.
Published: (2025)
RWKV-UI: UI Understanding with Enhanced Perception and Reasoning
by: Yang, Jiaxi, et al.
Published: (2025)
by: Yang, Jiaxi, et al.
Published: (2025)
Generative UI: LLMs are Effective UI Generators
by: Leviathan, Yaniv, et al.
Published: (2026)
by: Leviathan, Yaniv, et al.
Published: (2026)
Open WebUI: An Open, Extensible, and Usable Interface for AI Interaction
by: Baek, Jaeryang, et al.
Published: (2025)
by: Baek, Jaeryang, et al.
Published: (2025)
Similar Items
-
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
by: You, Keen, et al.
Published: (2024) -
Prioritising Interactive Flows in Data Center Networks With Central Control
by: Moorthy, Mohana Prasad Sathya
Published: (2023) -
Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
by: Yang, Zhen, et al.
Published: (2025) -
Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
by: Zhang, Haotian, et al.
Published: (2024) -
How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
by: Qian, Yusu, et al.
Published: (2024)