AutoDroid: LLM-powered Task Automation in Android

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wen, Hao, Li, Yuanchun, Liu, Guohong, Zhao, Shanhui, Yu, Tao, Li, Toby Jia-Jun, Jiang, Shiqi, Liu, Yunhao, Zhang, Yaqin, Liu, Yunxin
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917608815591424
author Wen, Hao
Li, Yuanchun
Liu, Guohong
Zhao, Shanhui
Yu, Tao
Li, Toby Jia-Jun
Jiang, Shiqi
Liu, Yunhao
Zhang, Yaqin
Liu, Yunxin
author_facet Wen, Hao
Li, Yuanchun
Liu, Guohong
Zhao, Shanhui
Yu, Tao
Li, Toby Jia-Jun
Jiang, Shiqi
Liu, Yunhao
Zhang, Yaqin
Liu, Yunxin
contents Mobile task automation is an attractive technique that aims to enable voice-based hands-free user interaction with smartphones. However, existing approaches suffer from poor scalability due to the limited language understanding ability and the non-trivial manual efforts required from developers or end-users. The recent advance of large language models (LLMs) in language understanding and reasoning inspires us to rethink the problem from a model-centric perspective, where task preparation, comprehension, and execution are handled by a unified language model. In this work, we introduce AutoDroid, a mobile task automation system capable of handling arbitrary tasks on any Android application without manual efforts. The key insight is to combine the commonsense knowledge of LLMs and domain-specific knowledge of apps through automated dynamic analysis. The main components include a functionality-aware UI representation method that bridges the UI with the LLM, exploration-based memory injection techniques that augment the app-specific domain knowledge of LLM, and a multi-granularity query optimization module that reduces the cost of model inference. We integrate AutoDroid with off-the-shelf LLMs including online GPT-4/GPT-3.5 and on-device Vicuna, and evaluate its performance on a new benchmark for memory-augmented Android task automation with 158 common tasks. The results demonstrated that AutoDroid is able to precisely generate actions with an accuracy of 90.9%, and complete tasks with a success rate of 71.3%, outperforming the GPT-4-powered baselines by 36.4% and 39.7%. The demo, benchmark suites, and source code of AutoDroid will be released at url{https://autodroid-sys.github.io/}.
format Preprint
id arxiv_https___arxiv_org_abs_2308_15272
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle AutoDroid: LLM-powered Task Automation in Android
Wen, Hao
Li, Yuanchun
Liu, Guohong
Zhao, Shanhui
Yu, Tao
Li, Toby Jia-Jun
Jiang, Shiqi
Liu, Yunhao
Zhang, Yaqin
Liu, Yunxin
Artificial Intelligence
Software Engineering
Mobile task automation is an attractive technique that aims to enable voice-based hands-free user interaction with smartphones. However, existing approaches suffer from poor scalability due to the limited language understanding ability and the non-trivial manual efforts required from developers or end-users. The recent advance of large language models (LLMs) in language understanding and reasoning inspires us to rethink the problem from a model-centric perspective, where task preparation, comprehension, and execution are handled by a unified language model. In this work, we introduce AutoDroid, a mobile task automation system capable of handling arbitrary tasks on any Android application without manual efforts. The key insight is to combine the commonsense knowledge of LLMs and domain-specific knowledge of apps through automated dynamic analysis. The main components include a functionality-aware UI representation method that bridges the UI with the LLM, exploration-based memory injection techniques that augment the app-specific domain knowledge of LLM, and a multi-granularity query optimization module that reduces the cost of model inference. We integrate AutoDroid with off-the-shelf LLMs including online GPT-4/GPT-3.5 and on-device Vicuna, and evaluate its performance on a new benchmark for memory-augmented Android task automation with 158 common tasks. The results demonstrated that AutoDroid is able to precisely generate actions with an accuracy of 90.9%, and complete tasks with a success rate of 71.3%, outperforming the GPT-4-powered baselines by 36.4% and 39.7%. The demo, benchmark suites, and source code of AutoDroid will be released at url{https://autodroid-sys.github.io/}.
title AutoDroid: LLM-powered Task Automation in Android
topic Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2308.15272