TinyClick: Single-Turn Agent for Empowering GUI Automation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915295315099648 |
|---|---|
| author | Pawlowski, Pawel Zawistowski, Krystian Lapacz, Wojciech Wiacek, Adam Skorupa, Marcin Postansque, Sebastien Hoscilowicz, Jakub |
| author_facet | Pawlowski, Pawel Zawistowski, Krystian Lapacz, Wojciech Wiacek, Adam Skorupa, Marcin Postansque, Sebastien Hoscilowicz, Jakub |
| contents | We present an UI agent for user interface (UI) interaction tasks, using Vision-Language Model Florence-2-Base. The agent's primary task is identifying the screen coordinates of the UI element corresponding to the user's command. It demonstrates very strong performance on Screenspot and OmniAct annotations, while maintaining a very small size of 0.27B parameters and minimal latency. Moreover, training needs small compute budget of 56 GPU-hours (worth about 40 USD). Relevant improvement comes from vision-specific multi-task training and MLLM-based data augmentation. We hope that decreased needs for expensive compute resources and manually annotated data will allow to facilitate more inclusive and sustainable research of UI agents. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_11871 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | TinyClick: Single-Turn Agent for Empowering GUI Automation Pawlowski, Pawel Zawistowski, Krystian Lapacz, Wojciech Wiacek, Adam Skorupa, Marcin Postansque, Sebastien Hoscilowicz, Jakub Human-Computer Interaction Artificial Intelligence We present an UI agent for user interface (UI) interaction tasks, using Vision-Language Model Florence-2-Base. The agent's primary task is identifying the screen coordinates of the UI element corresponding to the user's command. It demonstrates very strong performance on Screenspot and OmniAct annotations, while maintaining a very small size of 0.27B parameters and minimal latency. Moreover, training needs small compute budget of 56 GPU-hours (worth about 40 USD). Relevant improvement comes from vision-specific multi-task training and MLLM-based data augmentation. We hope that decreased needs for expensive compute resources and manually annotated data will allow to facilitate more inclusive and sustainable research of UI agents. |
| title | TinyClick: Single-Turn Agent for Empowering GUI Automation |
| topic | Human-Computer Interaction Artificial Intelligence |
| url | https://arxiv.org/abs/2410.11871 |