AppVLM: A Lightweight Vision Language Model for Online App Control

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Papoudakis, Georgios, Coste, Thomas, Wu, Zhihao, Hao, Jianye, Wang, Jun, Shao, Kun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912226332377088
author Papoudakis, Georgios
Coste, Thomas
Wu, Zhihao
Hao, Jianye
Wang, Jun
Shao, Kun
author_facet Papoudakis, Georgios
Coste, Thomas
Wu, Zhihao
Hao, Jianye
Wang, Jun
Shao, Kun
contents The utilisation of foundation models as smartphone assistants, termed app agents, is a critical research challenge. These agents aim to execute human instructions on smartphones by interpreting textual instructions and performing actions via the device's interface. While promising, current approaches face significant limitations. Methods that use large proprietary models, such as GPT-4o, are computationally expensive, while those that use smaller fine-tuned models often lack adaptability to out-of-distribution tasks. In this work, we introduce AppVLM, a lightweight Vision-Language Model (VLM). First, we fine-tune it offline on the AndroidControl dataset. Then, we refine its policy by collecting data from the AndroidWorld environment and performing further training iterations. Our results indicate that AppVLM achieves the highest action prediction accuracy in offline evaluation on the AndroidControl dataset, compared to all evaluated baselines, and matches GPT-4o in online task completion success rate in the AndroidWorld environment, while being up to ten times faster. This makes AppVLM a practical and efficient solution for real-world deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2502_06395
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AppVLM: A Lightweight Vision Language Model for Online App Control
Papoudakis, Georgios
Coste, Thomas
Wu, Zhihao
Hao, Jianye
Wang, Jun
Shao, Kun
Artificial Intelligence
The utilisation of foundation models as smartphone assistants, termed app agents, is a critical research challenge. These agents aim to execute human instructions on smartphones by interpreting textual instructions and performing actions via the device's interface. While promising, current approaches face significant limitations. Methods that use large proprietary models, such as GPT-4o, are computationally expensive, while those that use smaller fine-tuned models often lack adaptability to out-of-distribution tasks. In this work, we introduce AppVLM, a lightweight Vision-Language Model (VLM). First, we fine-tune it offline on the AndroidControl dataset. Then, we refine its policy by collecting data from the AndroidWorld environment and performing further training iterations. Our results indicate that AppVLM achieves the highest action prediction accuracy in offline evaluation on the AndroidControl dataset, compared to all evaluated baselines, and matches GPT-4o in online task completion success rate in the AndroidWorld environment, while being up to ten times faster. This makes AppVLM a practical and efficient solution for real-world deployment.
title AppVLM: A Lightweight Vision Language Model for Online App Control
topic Artificial Intelligence
url https://arxiv.org/abs/2502.06395