Harnessing Webpage UIs for Text-Rich Visual Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Junpeng, Ou, Tianyue, Song, Yifan, Qu, Yuxiao, Lam, Wai, Xiong, Chenyan, Chen, Wenhu, Neubig, Graham, Yue, Xiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909378833022976
author Liu, Junpeng
Ou, Tianyue
Song, Yifan
Qu, Yuxiao
Lam, Wai
Xiong, Chenyan
Chen, Wenhu
Neubig, Graham
Yue, Xiang
author_facet Liu, Junpeng
Ou, Tianyue
Song, Yifan
Qu, Yuxiao
Lam, Wai
Xiong, Chenyan
Chen, Wenhu
Neubig, Graham
Yue, Xiang
contents Text-rich visual understanding-the ability to process environments where dense textual content is integrated with visuals-is crucial for multimodal large language models (MLLMs) to interact effectively with structured environments. To enhance this capability, we propose synthesizing general multimodal instructions from webpage UIs using text-based large language models (LLMs). Despite lacking direct visual input, text-based LLMs are able to process structured text representations from webpage accessibility trees. These instructions are then paired with UI screenshots to train multimodal models. We introduce MultiUI, a dataset containing 7.3 million samples from 1 million websites, covering diverse multimodal tasks and UI layouts. Models trained on MultiUI not only excel in web UI tasks-achieving up to a 48% improvement on VisualWebBench and a 19.1% boost in element accuracy on a web agent dataset Mind2Web-but also generalize surprisingly well to non-web UI tasks and even to non-UI domains, such as document understanding, OCR, and chart interpretation. These results highlight the broad applicability of web UI data for advancing text-rich visual understanding across various scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13824
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Harnessing Webpage UIs for Text-Rich Visual Understanding
Liu, Junpeng
Ou, Tianyue
Song, Yifan
Qu, Yuxiao
Lam, Wai
Xiong, Chenyan
Chen, Wenhu
Neubig, Graham
Yue, Xiang
Computer Vision and Pattern Recognition
Computation and Language
Text-rich visual understanding-the ability to process environments where dense textual content is integrated with visuals-is crucial for multimodal large language models (MLLMs) to interact effectively with structured environments. To enhance this capability, we propose synthesizing general multimodal instructions from webpage UIs using text-based large language models (LLMs). Despite lacking direct visual input, text-based LLMs are able to process structured text representations from webpage accessibility trees. These instructions are then paired with UI screenshots to train multimodal models. We introduce MultiUI, a dataset containing 7.3 million samples from 1 million websites, covering diverse multimodal tasks and UI layouts. Models trained on MultiUI not only excel in web UI tasks-achieving up to a 48% improvement on VisualWebBench and a 19.1% boost in element accuracy on a web agent dataset Mind2Web-but also generalize surprisingly well to non-web UI tasks and even to non-UI domains, such as document understanding, OCR, and chart interpretation. These results highlight the broad applicability of web UI data for advancing text-rich visual understanding across various scenarios.
title Harnessing Webpage UIs for Text-Rich Visual Understanding
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2410.13824