InfoDet: A Dataset for Infographic Element Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Jiangning, Zhou, Yuxing, Wang, Zheng, Yao, Juntao, Gu, Yima, Yuan, Yuhui, Liu, Shixia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911212929810432
author Zhu, Jiangning
Zhou, Yuxing
Wang, Zheng
Yao, Juntao
Gu, Yima
Yuan, Yuhui
Liu, Shixia
author_facet Zhu, Jiangning
Zhou, Yuxing
Wang, Zheng
Yao, Juntao
Gu, Yima
Yuan, Yuhui
Liu, Shixia
contents Given the central role of charts in scientific, business, and communication contexts, enhancing the chart understanding capabilities of vision-language models (VLMs) has become increasingly critical. A key limitation of existing VLMs lies in their inaccurate visual grounding of infographic elements, including charts and human-recognizable objects (HROs) such as icons and images. However, chart understanding often requires identifying relevant elements and reasoning over them. To address this limitation, we introduce InfoDet, a dataset designed to support the development of accurate object detection models for charts and HROs in infographics. It contains 11,264 real and 90,000 synthetic infographics, with over 14 million bounding box annotations. These annotations are created by combining the model-in-the-loop and programmatic methods. We demonstrate the usefulness of InfoDet through three applications: 1) constructing a Thinking-with-Boxes scheme to boost the chart understanding performance of VLMs, 2) comparing existing object detection models, and 3) applying the developed detection model to document layout and UI element detection.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17473
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InfoDet: A Dataset for Infographic Element Detection
Zhu, Jiangning
Zhou, Yuxing
Wang, Zheng
Yao, Juntao
Gu, Yima
Yuan, Yuhui
Liu, Shixia
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Given the central role of charts in scientific, business, and communication contexts, enhancing the chart understanding capabilities of vision-language models (VLMs) has become increasingly critical. A key limitation of existing VLMs lies in their inaccurate visual grounding of infographic elements, including charts and human-recognizable objects (HROs) such as icons and images. However, chart understanding often requires identifying relevant elements and reasoning over them. To address this limitation, we introduce InfoDet, a dataset designed to support the development of accurate object detection models for charts and HROs in infographics. It contains 11,264 real and 90,000 synthetic infographics, with over 14 million bounding box annotations. These annotations are created by combining the model-in-the-loop and programmatic methods. We demonstrate the usefulness of InfoDet through three applications: 1) constructing a Thinking-with-Boxes scheme to boost the chart understanding performance of VLMs, 2) comparing existing object detection models, and 3) applying the developed detection model to document layout and UI element detection.
title InfoDet: A Dataset for Infographic Element Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.17473