InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Yuhang, Li, Pengxiang, Wei, Zishu, Xie, Congkai, Hu, Xueyu, Xu, Xinchen, Zhang, Shengyu, Han, Xiaotian, Yang, Hongxia, Wu, Fei
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916556875759616
author Liu, Yuhang
Li, Pengxiang
Wei, Zishu
Xie, Congkai
Hu, Xueyu
Xu, Xinchen
Zhang, Shengyu
Han, Xiaotian
Yang, Hongxia
Wu, Fei
author_facet Liu, Yuhang
Li, Pengxiang
Wei, Zishu
Xie, Congkai
Hu, Xueyu
Xu, Xinchen
Zhang, Shengyu
Han, Xiaotian
Yang, Hongxia
Wu, Fei
contents Graphical User Interface (GUI) Agents, powered by multimodal large language models (MLLMs), have shown great potential for task automation on computing devices such as computers and mobile phones. However, existing agents face challenges in multi-step reasoning and reliance on textual annotations, limiting their effectiveness. We introduce \textit{InfiGUIAgent}, an MLLM-based GUI Agent trained with a two-stage supervised fine-tuning pipeline. Stage 1 enhances fundamental skills such as GUI understanding and grounding, while Stage 2 integrates hierarchical reasoning and expectation-reflection reasoning skills using synthesized data to enable native reasoning abilities of the agents. \textit{InfiGUIAgent} achieves competitive performance on several GUI benchmarks, highlighting the impact of native reasoning skills in enhancing GUI interaction for automation tasks. Resources are available at \url{https://github.com/Reallm-Labs/InfiGUIAgent}.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04575
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
Liu, Yuhang
Li, Pengxiang
Wei, Zishu
Xie, Congkai
Hu, Xueyu
Xu, Xinchen
Zhang, Shengyu
Han, Xiaotian
Yang, Hongxia
Wu, Fei
Artificial Intelligence
Computation and Language
Human-Computer Interaction
Graphical User Interface (GUI) Agents, powered by multimodal large language models (MLLMs), have shown great potential for task automation on computing devices such as computers and mobile phones. However, existing agents face challenges in multi-step reasoning and reliance on textual annotations, limiting their effectiveness. We introduce \textit{InfiGUIAgent}, an MLLM-based GUI Agent trained with a two-stage supervised fine-tuning pipeline. Stage 1 enhances fundamental skills such as GUI understanding and grounding, while Stage 2 integrates hierarchical reasoning and expectation-reflection reasoning skills using synthesized data to enable native reasoning abilities of the agents. \textit{InfiGUIAgent} achieves competitive performance on several GUI benchmarks, highlighting the impact of native reasoning skills in enhancing GUI interaction for automation tasks. Resources are available at \url{https://github.com/Reallm-Labs/InfiGUIAgent}.
title InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
topic Artificial Intelligence
Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2501.04575