GEBench: Benchmarking Image Generation Models as GUI Environments
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Haodong, Wu, Jingwei, Sun, Quan, Li, Guopeng, Tian, Juanxi, Zhang, Huanyu, Lai, Yanlin, An, Ruichuan, Peng, Hongbo, Dai, Yuhong, Li, Chenxi, Qing, Chunmei, Wang, Jia, Meng, Ziyang, Ge, Zheng, Zhang, Xiangyu, Jiang, Daxin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
di: Li, Haodong, et al.
Pubblicazione: (2026)
di: Li, Haodong, et al.
Pubblicazione: (2026)
WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual Rubrics
di: Dai, Yuhong, et al.
Pubblicazione: (2026)
di: Dai, Yuhong, et al.
Pubblicazione: (2026)
PEARL: Personalized Streaming Video Understanding Model
di: Zheng, Yuanhong, et al.
Pubblicazione: (2026)
di: Zheng, Yuanhong, et al.
Pubblicazione: (2026)
GENIUS: Generative Fluid Intelligence Evaluation Suite
di: An, Ruichuan, et al.
Pubblicazione: (2026)
di: An, Ruichuan, et al.
Pubblicazione: (2026)
How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing
di: Zhang, Huanyu, et al.
Pubblicazione: (2026)
di: Zhang, Huanyu, et al.
Pubblicazione: (2026)
Benchmarking and Improving GUI Agents in High-Dynamic Environments
di: Liu, Enqi, et al.
Pubblicazione: (2026)
di: Liu, Enqi, et al.
Pubblicazione: (2026)
AutoGUI-v2: A Comprehensive Multi-Modal GUI Functionality Understanding Benchmark
di: Li, Hongxin, et al.
Pubblicazione: (2026)
di: Li, Hongxin, et al.
Pubblicazione: (2026)
GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
di: Yan, Haolong, et al.
Pubblicazione: (2025)
di: Yan, Haolong, et al.
Pubblicazione: (2025)
R-Align: Enhancing Generative Reward Models through Rationale-Centric Meta-Judging
di: Lai, Yanlin, et al.
Pubblicazione: (2026)
di: Lai, Yanlin, et al.
Pubblicazione: (2026)
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
di: Henry, Felix, et al.
Pubblicazione: (2026)
di: Henry, Felix, et al.
Pubblicazione: (2026)
SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Environments
di: Li, Dinging, et al.
Pubblicazione: (2026)
di: Li, Dinging, et al.
Pubblicazione: (2026)
WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments
di: Li, Jinchao, et al.
Pubblicazione: (2026)
di: Li, Jinchao, et al.
Pubblicazione: (2026)
Preclinical Evaluation of the Oral Toxicity, Genotoxicity, and Safety Pharmacology of LPM4870108, a Novel Potent Tropomyosin Receptor Kinase Inhibitor
di: Xiaochen Zhang, et al.
Pubblicazione: (2025)
di: Xiaochen Zhang, et al.
Pubblicazione: (2025)
Derived logarithmic deformation theory and moduli stacks of derived logarithmic structures
di: Zhang, Ruichuan
Pubblicazione: (2026)
di: Zhang, Ruichuan
Pubblicazione: (2026)
PRIME: A Process-Outcome Alignment Benchmark for Verifiable Reasoning in Mathematics and Engineering
di: Wang, Xiangfeng, et al.
Pubblicazione: (2026)
di: Wang, Xiangfeng, et al.
Pubblicazione: (2026)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
di: Ye, Xianhang, et al.
Pubblicazione: (2025)
di: Ye, Xianhang, et al.
Pubblicazione: (2025)
Deep-water and shallow-water limits of the intermediate long wave equation
di: Li, Guopeng
Pubblicazione: (2022)
di: Li, Guopeng
Pubblicazione: (2022)
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
di: Li, Hongxin, et al.
Pubblicazione: (2025)
di: Li, Hongxin, et al.
Pubblicazione: (2025)
EntWorld: A Holistic Environment and Benchmark for Verifiable Enterprise GUI Agents
di: Mo, Ying, et al.
Pubblicazione: (2026)
di: Mo, Ying, et al.
Pubblicazione: (2026)
Environmental Injection Attacks against GUI Agents in Realistic Dynamic Environments
di: Zhang, Yitong, et al.
Pubblicazione: (2025)
di: Zhang, Yitong, et al.
Pubblicazione: (2025)
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
di: Wang, Xuehui, et al.
Pubblicazione: (2025)
di: Wang, Xuehui, et al.
Pubblicazione: (2025)
Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
di: Tian, Juanxi, et al.
Pubblicazione: (2025)
di: Tian, Juanxi, et al.
Pubblicazione: (2025)
InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
di: Zhang, Ziyun, et al.
Pubblicazione: (2026)
di: Zhang, Ziyun, et al.
Pubblicazione: (2026)
MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments
di: Liu, Guangyi, et al.
Pubblicazione: (2026)
di: Liu, Guangyi, et al.
Pubblicazione: (2026)
Distributed Rotary Coverage Control of Multi-Agent Systems in Uncertain Environments
di: Zhai, Chao, et al.
Pubblicazione: (2025)
di: Zhai, Chao, et al.
Pubblicazione: (2025)
MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
di: Shi, Yucheng, et al.
Pubblicazione: (2025)
di: Shi, Yucheng, et al.
Pubblicazione: (2025)
DDFAV: Remote Sensing Large Vision Language Models Dataset and Evaluation Benchmark
di: Li, Haodong, et al.
Pubblicazione: (2024)
di: Li, Haodong, et al.
Pubblicazione: (2024)
Aria-UI: Visual Grounding for GUI Instructions
di: Yang, Yuhao, et al.
Pubblicazione: (2024)
di: Yang, Yuhao, et al.
Pubblicazione: (2024)
Energy Score-based Pseudo-Label Filtering and Adaptive Loss for Imbalanced Semi-supervised SAR target recognition
di: Zhang, Xinzheng, et al.
Pubblicazione: (2024)
di: Zhang, Xinzheng, et al.
Pubblicazione: (2024)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
di: Chen, Dongping, et al.
Pubblicazione: (2024)
di: Chen, Dongping, et al.
Pubblicazione: (2024)
UIPro: Unleashing Superior Interaction Capability For GUI Agents
di: Li, Hongxin, et al.
Pubblicazione: (2025)
di: Li, Hongxin, et al.
Pubblicazione: (2025)
SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI Environments
di: Ahmed, Syed Yusuf, et al.
Pubblicazione: (2026)
di: Ahmed, Syed Yusuf, et al.
Pubblicazione: (2026)
Global well-posedness of the 4-d energy-critical stochastic nonlinear Schrödinger equations with non-vanishing boundary condition
di: Cheung, Kelvin, et al.
Pubblicazione: (2019)
di: Cheung, Kelvin, et al.
Pubblicazione: (2019)
Persuasion in Networks With Strategic Substitutes
di: Guopeng Li, et al.
Pubblicazione: (2025)
di: Guopeng Li, et al.
Pubblicazione: (2025)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
Step-GUI Technical Report
di: Yan, Haolong, et al.
Pubblicazione: (2025)
di: Yan, Haolong, et al.
Pubblicazione: (2025)
GUI-GENESIS: Automated Synthesis of Efficient Environments with Verifiable Rewards for GUI Agent Post-Training
di: Cao, Yuan, et al.
Pubblicazione: (2026)
di: Cao, Yuan, et al.
Pubblicazione: (2026)
M-DocSum: Do LVLMs Genuinely Comprehend Interleaved Image-Text in Document Summarization?
di: Yan, Haolong, et al.
Pubblicazione: (2025)
di: Yan, Haolong, et al.
Pubblicazione: (2025)
Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria
di: Tian, Juanxi, et al.
Pubblicazione: (2026)
di: Tian, Juanxi, et al.
Pubblicazione: (2026)
MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
di: Chen, Ruihan, et al.
Pubblicazione: (2025)
di: Chen, Ruihan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
di: Li, Haodong, et al.
Pubblicazione: (2026) -
WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual Rubrics
di: Dai, Yuhong, et al.
Pubblicazione: (2026) -
PEARL: Personalized Streaming Video Understanding Model
di: Zheng, Yuanhong, et al.
Pubblicazione: (2026) -
GENIUS: Generative Fluid Intelligence Evaluation Suite
di: An, Ruichuan, et al.
Pubblicazione: (2026) -
How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing
di: Zhang, Huanyu, et al.
Pubblicazione: (2026)