OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Henry, Felix, Lin, Xiaochen, Zhu, Jiangyou, Yangfan, Zhang, Bingqian, Chen, Min, Huang, Shiyu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913143378149376
author Henry, Felix
Lin, Xiaochen
Zhu, Jiangyou
Yangfan
Zhang, Bingqian
Chen, Min
Huang, Shiyu
author_facet Henry, Felix
Lin, Xiaochen
Zhu, Jiangyou
Yangfan
Zhang, Bingqian
Chen, Min
Huang, Shiyu
contents Current benchmarks for graphical user interface (GUI) agents predominantly rely on static screenshots. However, real-world smartphone interaction routinely requires agents to process transient audio cues and temporal video dynamics that are tightly coupled with the moment of action. To bridge this gap, we introduce OmniGUI, the first step-level benchmark designed to evaluate GUI agents in omni-modal smartphone environments. OmniGUI provides continuous, interleaved multimodal inputs comprising static images, synchronous audio, and video clips at every action step. The dataset encompasses 709 expert-demonstrated episodes (2,579 action steps) across 29 applications, systematically annotated with objective multimodal dependency levels. Because dedicated omni-modal GUI agent frameworks are currently in their nascent stage, we select foundational omni-modal models capable of natively processing interleaved inputs to serve as agent proxies for our initial baselines. Our empirical evaluation reveals that while current models exhibit competency on visually static tasks, their action prediction performance degrades significantly in environments requiring synchronous temporal and auditory signals. Furthermore, ablation studies isolate specific operational bottlenecks, notably cross-modal interference when processing task-irrelevant environmental noise. The complete dataset, evaluation pipeline, and baseline prompts are provided in the supplementary material. Project page: https://omni-gui.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18758
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
Henry, Felix
Lin, Xiaochen
Zhu, Jiangyou
Yangfan
Zhang, Bingqian
Chen, Min
Huang, Shiyu
Human-Computer Interaction
Artificial Intelligence
Current benchmarks for graphical user interface (GUI) agents predominantly rely on static screenshots. However, real-world smartphone interaction routinely requires agents to process transient audio cues and temporal video dynamics that are tightly coupled with the moment of action. To bridge this gap, we introduce OmniGUI, the first step-level benchmark designed to evaluate GUI agents in omni-modal smartphone environments. OmniGUI provides continuous, interleaved multimodal inputs comprising static images, synchronous audio, and video clips at every action step. The dataset encompasses 709 expert-demonstrated episodes (2,579 action steps) across 29 applications, systematically annotated with objective multimodal dependency levels. Because dedicated omni-modal GUI agent frameworks are currently in their nascent stage, we select foundational omni-modal models capable of natively processing interleaved inputs to serve as agent proxies for our initial baselines. Our empirical evaluation reveals that while current models exhibit competency on visually static tasks, their action prediction performance degrades significantly in environments requiring synchronous temporal and auditory signals. Furthermore, ablation studies isolate specific operational bottlenecks, notably cross-modal interference when processing task-irrelevant environmental noise. The complete dataset, evaluation pipeline, and baseline prompts are provided in the supplementary material. Project page: https://omni-gui.github.io.
title OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2605.18758