TinyClick: Single-Turn Agent for Empowering GUI Automation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Pawlowski, Pawel, Zawistowski, Krystian, Lapacz, Wojciech, Wiacek, Adam, Skorupa, Marcin, Postansque, Sebastien, Hoscilowicz, Jakub
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915295315099648
author Pawlowski, Pawel
Zawistowski, Krystian
Lapacz, Wojciech
Wiacek, Adam
Skorupa, Marcin
Postansque, Sebastien
Hoscilowicz, Jakub
author_facet Pawlowski, Pawel
Zawistowski, Krystian
Lapacz, Wojciech
Wiacek, Adam
Skorupa, Marcin
Postansque, Sebastien
Hoscilowicz, Jakub
contents We present an UI agent for user interface (UI) interaction tasks, using Vision-Language Model Florence-2-Base. The agent's primary task is identifying the screen coordinates of the UI element corresponding to the user's command. It demonstrates very strong performance on Screenspot and OmniAct annotations, while maintaining a very small size of 0.27B parameters and minimal latency. Moreover, training needs small compute budget of 56 GPU-hours (worth about 40 USD). Relevant improvement comes from vision-specific multi-task training and MLLM-based data augmentation. We hope that decreased needs for expensive compute resources and manually annotated data will allow to facilitate more inclusive and sustainable research of UI agents.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11871
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TinyClick: Single-Turn Agent for Empowering GUI Automation
Pawlowski, Pawel
Zawistowski, Krystian
Lapacz, Wojciech
Wiacek, Adam
Skorupa, Marcin
Postansque, Sebastien
Hoscilowicz, Jakub
Human-Computer Interaction
Artificial Intelligence
We present an UI agent for user interface (UI) interaction tasks, using Vision-Language Model Florence-2-Base. The agent's primary task is identifying the screen coordinates of the UI element corresponding to the user's command. It demonstrates very strong performance on Screenspot and OmniAct annotations, while maintaining a very small size of 0.27B parameters and minimal latency. Moreover, training needs small compute budget of 56 GPU-hours (worth about 40 USD). Relevant improvement comes from vision-specific multi-task training and MLLM-based data augmentation. We hope that decreased needs for expensive compute resources and manually annotated data will allow to facilitate more inclusive and sustainable research of UI agents.
title TinyClick: Single-Turn Agent for Empowering GUI Automation
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2410.11871