Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2410.00201 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917792009158656 |
|---|---|
| author | Peng, Yi-Hao Huq, Faria Jiang, Yue Wu, Jason Li, Amanda Xin Yue Bigham, Jeffrey Pavel, Amy |
| author_facet | Peng, Yi-Hao Huq, Faria Jiang, Yue Wu, Jason Li, Amanda Xin Yue Bigham, Jeffrey Pavel, Amy |
| contents | Enabling machines to understand structured visuals like slides and user interfaces is essential for making them accessible to people with disabilities. However, achieving such understanding computationally has required manual data collection and annotation, which is time-consuming and labor-intensive. To overcome this challenge, we present a method to generate synthetic, structured visuals with target labels using code generation. Our method allows people to create datasets with built-in labels and train models with a small number of human-annotated examples. We demonstrate performance improvements in three tasks for understanding slides and UIs: recognizing visual elements, describing visual content, and classifying visual content types. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_00201 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | DreamStruct: Understanding Slides and User Interfaces via Synthetic Data Generation Peng, Yi-Hao Huq, Faria Jiang, Yue Wu, Jason Li, Amanda Xin Yue Bigham, Jeffrey Pavel, Amy Computer Vision and Pattern Recognition Computation and Language Enabling machines to understand structured visuals like slides and user interfaces is essential for making them accessible to people with disabilities. However, achieving such understanding computationally has required manual data collection and annotation, which is time-consuming and labor-intensive. To overcome this challenge, we present a method to generate synthetic, structured visuals with target labels using code generation. Our method allows people to create datasets with built-in labels and train models with a small number of human-annotated examples. We demonstrate performance improvements in three tasks for understanding slides and UIs: recognizing visual elements, describing visual content, and classifying visual content types. |
| title | DreamStruct: Understanding Slides and User Interfaces via Synthetic Data Generation |
| topic | Computer Vision and Pattern Recognition Computation and Language |
| url | https://arxiv.org/abs/2410.00201 |