SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jing, Hongyi, Chen, Jiafu, Rao, Chen, Dang, Ziqiang, Teng, Jiajie, Chu, Tianyi, Mo, Juncheng, Fang, Shuo, Lin, Huaizhong, Lv, Rui, Ma, Chenguang, Zhao, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912572153790464
author Jing, Hongyi
Chen, Jiafu
Rao, Chen
Dang, Ziqiang
Teng, Jiajie
Chu, Tianyi
Mo, Juncheng
Fang, Shuo
Lin, Huaizhong
Lv, Rui
Ma, Chenguang
Zhao, Lei
author_facet Jing, Hongyi
Chen, Jiafu
Rao, Chen
Dang, Ziqiang
Teng, Jiajie
Chu, Tianyi
Mo, Juncheng
Fang, Shuo
Lin, Huaizhong
Lv, Rui
Ma, Chenguang
Zhao, Lei
contents The existing Multimodal Large Language Models (MLLMs) for GUI perception have made great progress. However, the following challenges still exist in prior methods: 1) They model discrete coordinates based on text autoregressive mechanism, which results in lower grounding accuracy and slower inference speed. 2) They can only locate predefined sets of elements and are not capable of parsing the entire interface, which hampers the broad application and support for downstream tasks. To address the above issues, we propose SparkUI-Parser, a novel end-to-end framework where higher localization precision and fine-grained parsing capability of the entire interface are simultaneously achieved. Specifically, instead of using probability-based discrete modeling, we perform continuous modeling of coordinates based on a pre-trained Multimodal Large Language Model (MLLM) with an additional token router and coordinate decoder. This effectively mitigates the limitations inherent in the discrete output characteristics and the token-by-token generation process of MLLMs, consequently boosting both the accuracy and the inference speed. To further enhance robustness, a rejection mechanism based on a modified Hungarian matching algorithm is introduced, which empowers the model to identify and reject non-existent elements, thereby reducing false positives. Moreover, we present ScreenParse, a rigorously constructed benchmark to systematically assess structural perception capabilities of GUI models across diverse scenarios. Extensive experiments demonstrate that our approach consistently outperforms SOTA methods on ScreenSpot, ScreenSpot-v2, CAGUI-Grounding and ScreenParse benchmarks. The resources are available at https://github.com/antgroup/SparkUI-Parser.
format Preprint
id arxiv_https___arxiv_org_abs_2509_04908
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
Jing, Hongyi
Chen, Jiafu
Rao, Chen
Dang, Ziqiang
Teng, Jiajie
Chu, Tianyi
Mo, Juncheng
Fang, Shuo
Lin, Huaizhong
Lv, Rui
Ma, Chenguang
Zhao, Lei
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Human-Computer Interaction
The existing Multimodal Large Language Models (MLLMs) for GUI perception have made great progress. However, the following challenges still exist in prior methods: 1) They model discrete coordinates based on text autoregressive mechanism, which results in lower grounding accuracy and slower inference speed. 2) They can only locate predefined sets of elements and are not capable of parsing the entire interface, which hampers the broad application and support for downstream tasks. To address the above issues, we propose SparkUI-Parser, a novel end-to-end framework where higher localization precision and fine-grained parsing capability of the entire interface are simultaneously achieved. Specifically, instead of using probability-based discrete modeling, we perform continuous modeling of coordinates based on a pre-trained Multimodal Large Language Model (MLLM) with an additional token router and coordinate decoder. This effectively mitigates the limitations inherent in the discrete output characteristics and the token-by-token generation process of MLLMs, consequently boosting both the accuracy and the inference speed. To further enhance robustness, a rejection mechanism based on a modified Hungarian matching algorithm is introduced, which empowers the model to identify and reject non-existent elements, thereby reducing false positives. Moreover, we present ScreenParse, a rigorously constructed benchmark to systematically assess structural perception capabilities of GUI models across diverse scenarios. Extensive experiments demonstrate that our approach consistently outperforms SOTA methods on ScreenSpot, ScreenSpot-v2, CAGUI-Grounding and ScreenParse benchmarks. The resources are available at https://github.com/antgroup/SparkUI-Parser.
title SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Human-Computer Interaction
url https://arxiv.org/abs/2509.04908