VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917381237899264 |
|---|---|
| author | Wu, Zheng Huang, Heyuan Lou, Xingyu Qu, Xiangmou Cheng, Pengzhou Wu, Zongru Liu, Weiwen Zhang, Weinan Wang, Jun Wang, Zhaoxiang Zhang, Zhuosheng |
| author_facet | Wu, Zheng Huang, Heyuan Lou, Xingyu Qu, Xiangmou Cheng, Pengzhou Wu, Zongru Liu, Weiwen Zhang, Weinan Wang, Jun Wang, Zhaoxiang Zhang, Zhuosheng |
| contents | With the rapid progress of multimodal large language models, operating system (OS) agents become increasingly capable of automating tasks through on-device graphical user interfaces (GUIs). However, most existing OS agents are designed for idealized settings, whereas real-world environments often present untrustworthy conditions. To mitigate risks of over-execution in such scenarios, we propose a query-driven human-agent-GUI interaction framework that enables OS agents to decide when to query humans for more reliable task completion. Built upon this framework, we introduce VeriOS-Agent, a trustworthy OS agent trained with a three-stage learning paradigm that falicitate the decoupling and utilization of meta-knowledge by supervised fine-tuning and group relative policy optimization. Concretely, VeriOS-Agent autonomously executes actions in normal conditions while proactively querying humans in untrustworthy scenarios. Experiments show that VeriOS-Agent improves the average step-wise success rate by 19.72\% in over the strongest baselines, without compromising normal performance. VeriOS-Agent significantly improves performance in untrustworthy scenarios while maintaining comparable performance in trustworthy scenarios. Analysis highlights VeriOS-Agent's rationality, generalizability, and scalability. The codes, datasets and models are available at https://github.com/Wuzheng02/VeriOS. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_07553 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents Wu, Zheng Huang, Heyuan Lou, Xingyu Qu, Xiangmou Cheng, Pengzhou Wu, Zongru Liu, Weiwen Zhang, Weinan Wang, Jun Wang, Zhaoxiang Zhang, Zhuosheng Computation and Language With the rapid progress of multimodal large language models, operating system (OS) agents become increasingly capable of automating tasks through on-device graphical user interfaces (GUIs). However, most existing OS agents are designed for idealized settings, whereas real-world environments often present untrustworthy conditions. To mitigate risks of over-execution in such scenarios, we propose a query-driven human-agent-GUI interaction framework that enables OS agents to decide when to query humans for more reliable task completion. Built upon this framework, we introduce VeriOS-Agent, a trustworthy OS agent trained with a three-stage learning paradigm that falicitate the decoupling and utilization of meta-knowledge by supervised fine-tuning and group relative policy optimization. Concretely, VeriOS-Agent autonomously executes actions in normal conditions while proactively querying humans in untrustworthy scenarios. Experiments show that VeriOS-Agent improves the average step-wise success rate by 19.72\% in over the strongest baselines, without compromising normal performance. VeriOS-Agent significantly improves performance in untrustworthy scenarios while maintaining comparable performance in trustworthy scenarios. Analysis highlights VeriOS-Agent's rationality, generalizability, and scalability. The codes, datasets and models are available at https://github.com/Wuzheng02/VeriOS. |
| title | VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2509.07553 |