Saved in:
Bibliographic Details
Main Authors: Peng, Yi-Hao, Huq, Faria, Jiang, Yue, Wu, Jason, Li, Amanda Xin Yue, Bigham, Jeffrey, Pavel, Amy
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2410.00201
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917792009158656
author Peng, Yi-Hao
Huq, Faria
Jiang, Yue
Wu, Jason
Li, Amanda Xin Yue
Bigham, Jeffrey
Pavel, Amy
author_facet Peng, Yi-Hao
Huq, Faria
Jiang, Yue
Wu, Jason
Li, Amanda Xin Yue
Bigham, Jeffrey
Pavel, Amy
contents Enabling machines to understand structured visuals like slides and user interfaces is essential for making them accessible to people with disabilities. However, achieving such understanding computationally has required manual data collection and annotation, which is time-consuming and labor-intensive. To overcome this challenge, we present a method to generate synthetic, structured visuals with target labels using code generation. Our method allows people to create datasets with built-in labels and train models with a small number of human-annotated examples. We demonstrate performance improvements in three tasks for understanding slides and UIs: recognizing visual elements, describing visual content, and classifying visual content types.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00201
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DreamStruct: Understanding Slides and User Interfaces via Synthetic Data Generation
Peng, Yi-Hao
Huq, Faria
Jiang, Yue
Wu, Jason
Li, Amanda Xin Yue
Bigham, Jeffrey
Pavel, Amy
Computer Vision and Pattern Recognition
Computation and Language
Enabling machines to understand structured visuals like slides and user interfaces is essential for making them accessible to people with disabilities. However, achieving such understanding computationally has required manual data collection and annotation, which is time-consuming and labor-intensive. To overcome this challenge, we present a method to generate synthetic, structured visuals with target labels using code generation. Our method allows people to create datasets with built-in labels and train models with a small number of human-annotated examples. We demonstrate performance improvements in three tasks for understanding slides and UIs: recognizing visual elements, describing visual content, and classifying visual content types.
title DreamStruct: Understanding Slides and User Interfaces via Synthetic Data Generation
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2410.00201