UI-UG: A Unified MLLM for UI Understanding and Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yang, Hao, Qiu, Weijie, Zhang, Ru, Fang, Zhou, Mao, Ruichao, Lin, Xiaoyu, Huang, Maji, Huang, Zhaosong, Guo, Teng, Liu, Shuoyang, Rao, Hai
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912616813690880
author Yang, Hao
Qiu, Weijie
Zhang, Ru
Fang, Zhou
Mao, Ruichao
Lin, Xiaoyu
Huang, Maji
Huang, Zhaosong
Guo, Teng
Liu, Shuoyang
Rao, Hai
author_facet Yang, Hao
Qiu, Weijie
Zhang, Ru
Fang, Zhou
Mao, Ruichao
Lin, Xiaoyu
Huang, Maji
Huang, Zhaosong
Guo, Teng
Liu, Shuoyang
Rao, Hai
contents Although Multimodal Large Language Models (MLLMs) have been widely applied across domains, they are still facing challenges in domain-specific tasks, such as User Interface (UI) understanding accuracy and UI generation quality. In this paper, we introduce UI-UG (a unified MLLM for UI Understanding and Generation), integrating both capabilities. For understanding tasks, we employ Supervised Fine-tuning (SFT) combined with Group Relative Policy Optimization (GRPO) to enhance fine-grained understanding on the modern complex UI data. For generation tasks, we further use Direct Preference Optimization (DPO) to make our model generate human-preferred UIs. In addition, we propose an industrially effective workflow, including the design of an LLM-friendly domain-specific language (DSL), training strategies, rendering processes, and evaluation metrics. In experiments, our model achieves state-of-the-art (SOTA) performance on understanding tasks, outperforming both larger general-purpose MLLMs and similarly-sized UI-specialized models. Our model is also on par with these larger MLLMs in UI generation performance at a fraction of the computational cost. We also demonstrate that integrating understanding and generation tasks can improve accuracy and quality for both tasks. Code and Model: https://github.com/neovateai/UI-UG
format Preprint
id arxiv_https___arxiv_org_abs_2509_24361
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UI-UG: A Unified MLLM for UI Understanding and Generation
Yang, Hao
Qiu, Weijie
Zhang, Ru
Fang, Zhou
Mao, Ruichao
Lin, Xiaoyu
Huang, Maji
Huang, Zhaosong
Guo, Teng
Liu, Shuoyang
Rao, Hai
Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
Although Multimodal Large Language Models (MLLMs) have been widely applied across domains, they are still facing challenges in domain-specific tasks, such as User Interface (UI) understanding accuracy and UI generation quality. In this paper, we introduce UI-UG (a unified MLLM for UI Understanding and Generation), integrating both capabilities. For understanding tasks, we employ Supervised Fine-tuning (SFT) combined with Group Relative Policy Optimization (GRPO) to enhance fine-grained understanding on the modern complex UI data. For generation tasks, we further use Direct Preference Optimization (DPO) to make our model generate human-preferred UIs. In addition, we propose an industrially effective workflow, including the design of an LLM-friendly domain-specific language (DSL), training strategies, rendering processes, and evaluation metrics. In experiments, our model achieves state-of-the-art (SOTA) performance on understanding tasks, outperforming both larger general-purpose MLLMs and similarly-sized UI-specialized models. Our model is also on par with these larger MLLMs in UI generation performance at a fraction of the computational cost. We also demonstrate that integrating understanding and generation tasks can improve accuracy and quality for both tasks. Code and Model: https://github.com/neovateai/UI-UG
title UI-UG: A Unified MLLM for UI Understanding and Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2509.24361