AppAgent v2: Advanced Agent for Flexible Mobile Interactions

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Yanda, Zhang, Chi, Jiang, Wenjia, Yang, Wanqi, Fu, Bin, Cheng, Pei, Chen, Xin, Chen, Ling, Wei, Yunchao
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914041930186752
author Li, Yanda
Zhang, Chi
Jiang, Wenjia
Yang, Wanqi
Fu, Bin
Cheng, Pei
Chen, Xin
Chen, Ling
Wei, Yunchao
author_facet Li, Yanda
Zhang, Chi
Jiang, Wenjia
Yang, Wanqi
Fu, Bin
Cheng, Pei
Chen, Xin
Chen, Ling
Wei, Yunchao
contents With the advancement of Multimodal Large Language Models (MLLM), LLM-driven visual agents are increasingly impacting software interfaces, particularly those with graphical user interfaces. This work introduces a novel LLM-based multimodal agent framework for mobile devices. This framework, capable of navigating mobile devices, emulates human-like interactions. Our agent constructs a flexible action space that enhances adaptability across various applications including parser, text and vision descriptions. The agent operates through two main phases: exploration and deployment. During the exploration phase, functionalities of user interface elements are documented either through agent-driven or manual explorations into a customized structured knowledge base. In the deployment phase, RAG technology enables efficient retrieval and update from this knowledge base, thereby empowering the agent to perform tasks effectively and accurately. This includes performing complex, multi-step operations across various applications, thereby demonstrating the framework's adaptability and precision in handling customized task workflows. Our experimental results across various benchmarks demonstrate the framework's superior performance, confirming its effectiveness in real-world scenarios. Our code will be open source soon.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11824
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AppAgent v2: Advanced Agent for Flexible Mobile Interactions
Li, Yanda
Zhang, Chi
Jiang, Wenjia
Yang, Wanqi
Fu, Bin
Cheng, Pei
Chen, Xin
Chen, Ling
Wei, Yunchao
Human-Computer Interaction
Artificial Intelligence
With the advancement of Multimodal Large Language Models (MLLM), LLM-driven visual agents are increasingly impacting software interfaces, particularly those with graphical user interfaces. This work introduces a novel LLM-based multimodal agent framework for mobile devices. This framework, capable of navigating mobile devices, emulates human-like interactions. Our agent constructs a flexible action space that enhances adaptability across various applications including parser, text and vision descriptions. The agent operates through two main phases: exploration and deployment. During the exploration phase, functionalities of user interface elements are documented either through agent-driven or manual explorations into a customized structured knowledge base. In the deployment phase, RAG technology enables efficient retrieval and update from this knowledge base, thereby empowering the agent to perform tasks effectively and accurately. This includes performing complex, multi-step operations across various applications, thereby demonstrating the framework's adaptability and precision in handling customized task workflows. Our experimental results across various benchmarks demonstrate the framework's superior performance, confirming its effectiveness in real-world scenarios. Our code will be open source soon.
title AppAgent v2: Advanced Agent for Flexible Mobile Interactions
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2408.11824