GEBench: Benchmarking Image Generation Models as GUI Environments

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Haodong, Wu, Jingwei, Sun, Quan, Li, Guopeng, Tian, Juanxi, Zhang, Huanyu, Lai, Yanlin, An, Ruichuan, Peng, Hongbo, Dai, Yuhong, Li, Chenxi, Qing, Chunmei, Wang, Jia, Meng, Ziyang, Ge, Zheng, Zhang, Xiangyu, Jiang, Daxin
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910017407418368
author Li, Haodong
Wu, Jingwei
Sun, Quan
Li, Guopeng
Tian, Juanxi
Zhang, Huanyu
Lai, Yanlin
An, Ruichuan
Peng, Hongbo
Dai, Yuhong
Li, Chenxi
Qing, Chunmei
Wang, Jia
Meng, Ziyang
Ge, Zheng
Zhang, Xiangyu
Jiang, Daxin
author_facet Li, Haodong
Wu, Jingwei
Sun, Quan
Li, Guopeng
Tian, Juanxi
Zhang, Huanyu
Lai, Yanlin
An, Ruichuan
Peng, Hongbo
Dai, Yuhong
Li, Chenxi
Qing, Chunmei
Wang, Jia
Meng, Ziyang
Ge, Zheng
Zhang, Xiangyu
Jiang, Daxin
contents Recent advancements in image generation models have enabled the prediction of future Graphical User Interface (GUI) states based on user instructions. However, existing benchmarks primarily focus on general domain visual fidelity, leaving the evaluation of state transitions and temporal coherence in GUI-specific contexts underexplored. To address this gap, we introduce GEBench, a comprehensive benchmark for evaluating dynamic interaction and temporal coherence in GUI generation. GEBench comprises 700 carefully curated samples spanning five task categories, covering both single-step interactions and multi-step trajectories across real-world and fictional scenarios, as well as grounding point localization. To support systematic evaluation, we propose GE-Score, a novel five-dimensional metric that assesses Goal Achievement, Interaction Logic, Content Consistency, UI Plausibility, and Visual Quality. Extensive evaluations on current models indicate that while they perform well on single-step transitions, they struggle significantly with maintaining temporal coherence and spatial grounding over longer interaction sequences. Our findings identify icon interpretation, text rendering, and localization precision as critical bottlenecks. This work provides a foundation for systematic assessment and suggests promising directions for future research toward building high-fidelity generative GUI environments. The code is available at: https://github.com/stepfun-ai/GEBench.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09007
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GEBench: Benchmarking Image Generation Models as GUI Environments
Li, Haodong
Wu, Jingwei
Sun, Quan
Li, Guopeng
Tian, Juanxi
Zhang, Huanyu
Lai, Yanlin
An, Ruichuan
Peng, Hongbo
Dai, Yuhong
Li, Chenxi
Qing, Chunmei
Wang, Jia
Meng, Ziyang
Ge, Zheng
Zhang, Xiangyu
Jiang, Daxin
Artificial Intelligence
Computer Vision and Pattern Recognition
Recent advancements in image generation models have enabled the prediction of future Graphical User Interface (GUI) states based on user instructions. However, existing benchmarks primarily focus on general domain visual fidelity, leaving the evaluation of state transitions and temporal coherence in GUI-specific contexts underexplored. To address this gap, we introduce GEBench, a comprehensive benchmark for evaluating dynamic interaction and temporal coherence in GUI generation. GEBench comprises 700 carefully curated samples spanning five task categories, covering both single-step interactions and multi-step trajectories across real-world and fictional scenarios, as well as grounding point localization. To support systematic evaluation, we propose GE-Score, a novel five-dimensional metric that assesses Goal Achievement, Interaction Logic, Content Consistency, UI Plausibility, and Visual Quality. Extensive evaluations on current models indicate that while they perform well on single-step transitions, they struggle significantly with maintaining temporal coherence and spatial grounding over longer interaction sequences. Our findings identify icon interpretation, text rendering, and localization precision as critical bottlenecks. This work provides a foundation for systematic assessment and suggests promising directions for future research toward building high-fidelity generative GUI environments. The code is available at: https://github.com/stepfun-ai/GEBench.
title GEBench: Benchmarking Image Generation Models as GUI Environments
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.09007