Aria-UI: Visual Grounding for GUI Instructions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yuhao, Wang, Yue, Li, Dongxu, Luo, Ziyang, Chen, Bei, Huang, Chao, Li, Junnan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911043583737856
author Yang, Yuhao
Wang, Yue
Li, Dongxu
Luo, Ziyang
Chen, Bei
Huang, Chao
Li, Junnan
author_facet Yang, Yuhao
Wang, Yue
Li, Dongxu
Luo, Ziyang
Chen, Bei
Huang, Chao
Li, Junnan
contents Digital agents for automating tasks across different platforms by directly manipulating the GUIs are increasingly important. For these agents, grounding from language instructions to target elements remains a significant challenge due to reliance on HTML or AXTree inputs. In this paper, we introduce Aria-UI, a large multimodal model specifically designed for GUI grounding. Aria-UI adopts a pure-vision approach, eschewing reliance on auxiliary inputs. To adapt to heterogeneous planning instructions, we propose a scalable data pipeline that synthesizes diverse and high-quality instruction samples for grounding. To handle dynamic contexts in task performing, Aria-UI incorporates textual and text-image interleaved action histories, enabling robust context-aware reasoning for grounding. Aria-UI sets new state-of-the-art results across offline and online agent benchmarks, outperforming both vision-only and AXTree-reliant baselines. We release all training data and model checkpoints to foster further research at https://ariaui.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16256
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Aria-UI: Visual Grounding for GUI Instructions
Yang, Yuhao
Wang, Yue
Li, Dongxu
Luo, Ziyang
Chen, Bei
Huang, Chao
Li, Junnan
Human-Computer Interaction
Artificial Intelligence
Digital agents for automating tasks across different platforms by directly manipulating the GUIs are increasingly important. For these agents, grounding from language instructions to target elements remains a significant challenge due to reliance on HTML or AXTree inputs. In this paper, we introduce Aria-UI, a large multimodal model specifically designed for GUI grounding. Aria-UI adopts a pure-vision approach, eschewing reliance on auxiliary inputs. To adapt to heterogeneous planning instructions, we propose a scalable data pipeline that synthesizes diverse and high-quality instruction samples for grounding. To handle dynamic contexts in task performing, Aria-UI incorporates textual and text-image interleaved action histories, enabling robust context-aware reasoning for grounding. Aria-UI sets new state-of-the-art results across offline and online agent benchmarks, outperforming both vision-only and AXTree-reliant baselines. We release all training data and model checkpoints to foster further research at https://ariaui.github.io.
title Aria-UI: Visual Grounding for GUI Instructions
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2412.16256