OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Jiancong, Wang, Wenjin, Zhang, Zhuomeng, Liu, Zihan, Liu, Qi, Feng, Ke, Sun, Zixun, Yang, Yuedong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914069165899776
author Xie, Jiancong
Wang, Wenjin
Zhang, Zhuomeng
Liu, Zihan
Liu, Qi
Feng, Ke
Sun, Zixun
Yang, Yuedong
author_facet Xie, Jiancong
Wang, Wenjin
Zhang, Zhuomeng
Liu, Zihan
Liu, Qi
Feng, Ke
Sun, Zixun
Yang, Yuedong
contents Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities. However, evaluating their capacity for human-like understanding in One-Image Guides remains insufficiently explored. One-Image Guides are a visual format combining text, imagery, and symbols to present reorganized and structured information for easier comprehension, which are specifically designed for human viewing and inherently embody the characteristics of human perception and understanding. Here, we present OIG-Bench, a comprehensive benchmark focused on One-Image Guide understanding across diverse domains. To reduce the cost of manual annotation, we developed a semi-automated annotation pipeline in which multiple intelligent agents collaborate to generate preliminary image descriptions, assisting humans in constructing image-text pairs. With OIG-Bench, we have conducted a comprehensive evaluation of 29 state-of-the-art MLLMs, including both proprietary and open-source models. The results show that Qwen2.5-VL-72B performs the best among the evaluated models, with an overall accuracy of 77%. Nevertheless, all models exhibit notable weaknesses in semantic understanding and logical reasoning, indicating that current MLLMs still struggle to accurately interpret complex visual-text relationships. In addition, we also demonstrate that the proposed multi-agent annotation system outperforms all MLLMs in image captioning, highlighting its potential as both a high-quality image description generator and a valuable tool for future dataset construction. Datasets are available at https://github.com/XiejcSYSU/OIG-Bench.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00069
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding
Xie, Jiancong
Wang, Wenjin
Zhang, Zhuomeng
Liu, Zihan
Liu, Qi
Feng, Ke
Sun, Zixun
Yang, Yuedong
Computer Vision and Pattern Recognition
Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities. However, evaluating their capacity for human-like understanding in One-Image Guides remains insufficiently explored. One-Image Guides are a visual format combining text, imagery, and symbols to present reorganized and structured information for easier comprehension, which are specifically designed for human viewing and inherently embody the characteristics of human perception and understanding. Here, we present OIG-Bench, a comprehensive benchmark focused on One-Image Guide understanding across diverse domains. To reduce the cost of manual annotation, we developed a semi-automated annotation pipeline in which multiple intelligent agents collaborate to generate preliminary image descriptions, assisting humans in constructing image-text pairs. With OIG-Bench, we have conducted a comprehensive evaluation of 29 state-of-the-art MLLMs, including both proprietary and open-source models. The results show that Qwen2.5-VL-72B performs the best among the evaluated models, with an overall accuracy of 77%. Nevertheless, all models exhibit notable weaknesses in semantic understanding and logical reasoning, indicating that current MLLMs still struggle to accurately interpret complex visual-text relationships. In addition, we also demonstrate that the proposed multi-agent annotation system outperforms all MLLMs in image captioning, highlighting its potential as both a high-quality image description generator and a valuable tool for future dataset construction. Datasets are available at https://github.com/XiejcSYSU/OIG-Bench.
title OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.00069