LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909907127631872 |
|---|---|
| author | Liu, Guangyi Zhao, Pengxiang Liang, Yaozhen Liu, Liang Guo, Yaxuan Xiao, Han Lin, Weifeng Chai, Yuxiang Han, Yue Ren, Shuai Wang, Hao Liang, Xiaoyu Wang, WenHao Wu, Tianze Lu, Zhengxi Chen, Siheng LiLinghao Wang, Hao Xiong, Guanjing Liu, Yong Li, Hongsheng |
| author_facet | Liu, Guangyi Zhao, Pengxiang Liang, Yaozhen Liu, Liang Guo, Yaxuan Xiao, Han Lin, Weifeng Chai, Yuxiang Han, Yue Ren, Shuai Wang, Hao Liang, Xiaoyu Wang, WenHao Wu, Tianze Lu, Zhengxi Chen, Siheng LiLinghao Wang, Hao Xiong, Guanjing Liu, Yong Li, Hongsheng |
| contents | With the rapid rise of large language models (LLMs), phone automation has undergone transformative changes. This paper systematically reviews LLM-driven phone GUI agents, highlighting their evolution from script-based automation to intelligent, adaptive systems. We first contextualize key challenges, (i) limited generality, (ii) high maintenance overhead, and (iii) weak intent comprehension, and show how LLMs address these issues through advanced language understanding, multimodal perception, and robust decision-making. We then propose a taxonomy covering fundamental agent frameworks (single-agent, multi-agent, plan-then-act), modeling approaches (prompt engineering, training-based), and essential datasets and benchmarks. Furthermore, we detail task-specific architectures, supervised fine-tuning, and reinforcement learning strategies that bridge user intent and GUI operations. Finally, we discuss open challenges such as dataset diversity, on-device deployment efficiency, user-centric adaptation, and security concerns, offering forward-looking insights into this rapidly evolving field. By providing a structured overview and identifying pressing research gaps, this paper serves as a definitive reference for researchers and practitioners seeking to harness LLMs in designing scalable, user-friendly phone GUI agents. The collection of papers reviewed in this survey will be hosted and regularly updated on the GitHub repository: https://github.com/PhoneLLM/Awesome-LLM-Powered-Phone-GUI-Agents |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_19838 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects Liu, Guangyi Zhao, Pengxiang Liang, Yaozhen Liu, Liang Guo, Yaxuan Xiao, Han Lin, Weifeng Chai, Yuxiang Han, Yue Ren, Shuai Wang, Hao Liang, Xiaoyu Wang, WenHao Wu, Tianze Lu, Zhengxi Chen, Siheng LiLinghao Wang, Hao Xiong, Guanjing Liu, Yong Li, Hongsheng Human-Computer Interaction With the rapid rise of large language models (LLMs), phone automation has undergone transformative changes. This paper systematically reviews LLM-driven phone GUI agents, highlighting their evolution from script-based automation to intelligent, adaptive systems. We first contextualize key challenges, (i) limited generality, (ii) high maintenance overhead, and (iii) weak intent comprehension, and show how LLMs address these issues through advanced language understanding, multimodal perception, and robust decision-making. We then propose a taxonomy covering fundamental agent frameworks (single-agent, multi-agent, plan-then-act), modeling approaches (prompt engineering, training-based), and essential datasets and benchmarks. Furthermore, we detail task-specific architectures, supervised fine-tuning, and reinforcement learning strategies that bridge user intent and GUI operations. Finally, we discuss open challenges such as dataset diversity, on-device deployment efficiency, user-centric adaptation, and security concerns, offering forward-looking insights into this rapidly evolving field. By providing a structured overview and identifying pressing research gaps, this paper serves as a definitive reference for researchers and practitioners seeking to harness LLMs in designing scalable, user-friendly phone GUI agents. The collection of papers reviewed in this survey will be hosted and regularly updated on the GitHub repository: https://github.com/PhoneLLM/Awesome-LLM-Powered-Phone-GUI-Agents |
| title | LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects |
| topic | Human-Computer Interaction |
| url | https://arxiv.org/abs/2504.19838 |