See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Zongru, Mao, Rui, Tian, Zhiyuan, Cheng, Pengzhou, Ju, Tianjie, Wu, Zheng, Dong, Lingzhong, Sheng, Haiyue, Zhang, Zhuosheng, Liu, Gongshen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908895108136960
author Wu, Zongru
Mao, Rui
Tian, Zhiyuan
Cheng, Pengzhou
Ju, Tianjie
Wu, Zheng
Dong, Lingzhong
Sheng, Haiyue
Zhang, Zhuosheng
Liu, Gongshen
author_facet Wu, Zongru
Mao, Rui
Tian, Zhiyuan
Cheng, Pengzhou
Ju, Tianjie
Wu, Zheng
Dong, Lingzhong
Sheng, Haiyue
Zhang, Zhuosheng
Liu, Gongshen
contents The advent of multimodal agents facilitates effective interaction within graphical user interface (GUI), especially in ubiquitous GUI control. However, their inability to reliably execute toggle control instructions remains a key bottleneck. To investigate this, we construct a state control benchmark with binary toggle instructions derived from public datasets. Evaluation results of existing agents demonstrate their notable unreliability, particularly when the current toggle state already matches the desired state. To address the challenge, we propose State-aware Reasoning (StaR), a multimodal reasoning method that enables agents to perceive the current toggle state, infer the desired state from the instruction, and act accordingly. Experiments on four multimodal agents demonstrate that StaR can improve toggle instruction execution accuracy by over 30\%. Further evaluations on three public agentic benchmarks show that StaR also enhances general agentic task performance. Finally, evaluations on a dynamic environment highlight the potential of StaR for real-world applications. Code and benchmark: https://github.com/ZrW00/StaR.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13615
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
Wu, Zongru
Mao, Rui
Tian, Zhiyuan
Cheng, Pengzhou
Ju, Tianjie
Wu, Zheng
Dong, Lingzhong
Sheng, Haiyue
Zhang, Zhuosheng
Liu, Gongshen
Artificial Intelligence
Computation and Language
Human-Computer Interaction
The advent of multimodal agents facilitates effective interaction within graphical user interface (GUI), especially in ubiquitous GUI control. However, their inability to reliably execute toggle control instructions remains a key bottleneck. To investigate this, we construct a state control benchmark with binary toggle instructions derived from public datasets. Evaluation results of existing agents demonstrate their notable unreliability, particularly when the current toggle state already matches the desired state. To address the challenge, we propose State-aware Reasoning (StaR), a multimodal reasoning method that enables agents to perceive the current toggle state, infer the desired state from the instruction, and act accordingly. Experiments on four multimodal agents demonstrate that StaR can improve toggle instruction execution accuracy by over 30\%. Further evaluations on three public agentic benchmarks show that StaR also enhances general agentic task performance. Finally, evaluations on a dynamic environment highlight the potential of StaR for real-world applications. Code and benchmark: https://github.com/ZrW00/StaR.
title See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
topic Artificial Intelligence
Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2509.13615