Saved in:
Bibliographic Details
Main Authors: Zhang, Jianyi, Ji, Xu, Zhou, Ziyin, Zhou, Yuchen, Shi, Shubo, Wu, Haoyu, Li, Zhen, Liu, Shizhao
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.00323
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913969502945280
author Zhang, Jianyi
Ji, Xu
Zhou, Ziyin
Zhou, Yuchen
Shi, Shubo
Wu, Haoyu
Li, Zhen
Liu, Shizhao
author_facet Zhang, Jianyi
Ji, Xu
Zhou, Ziyin
Zhou, Yuchen
Shi, Shubo
Wu, Haoyu
Li, Zhen
Liu, Shizhao
contents Evaluating the performance of visual language models (VLMs) in graphic reasoning tasks has become an important research topic. However, VLMs still show obvious deficiencies in simulating human-level graphic reasoning capabilities, especially in complex graphic reasoning and abstract problem solving, which are less studied and existing studies only focus on simple graphics. To evaluate the performance of VLMs in complex graphic reasoning, we propose ReasonBench, the first evaluation benchmark focused on structured graphic reasoning tasks, which includes 1,613 questions from real-world intelligence tests. ReasonBench covers reasoning dimensions related to location, attribute, quantity, and multi-element tasks, providing a comprehensive evaluation of the performance of VLMs in spatial, relational, and abstract reasoning capabilities. We benchmark 11 mainstream VLMs (including closed-source and open-source models) and reveal significant limitations of current models. Based on these findings, we propose a dual optimization strategy: Diagrammatic Reasoning Chain (DiaCoT) enhances the interpretability of reasoning by decomposing layers, and ReasonTune enhances the task adaptability of model reasoning through training, all of which improves VLM performance by 33.5\%. All experimental data and code are in the repository: https://huggingface.co/datasets/cistine/ReasonBench.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00323
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Oedipus and the Sphinx: Benchmarking and Improving Visual Language Models for Complex Graphic Reasoning
Zhang, Jianyi
Ji, Xu
Zhou, Ziyin
Zhou, Yuchen
Shi, Shubo
Wu, Haoyu
Li, Zhen
Liu, Shizhao
Artificial Intelligence
Evaluating the performance of visual language models (VLMs) in graphic reasoning tasks has become an important research topic. However, VLMs still show obvious deficiencies in simulating human-level graphic reasoning capabilities, especially in complex graphic reasoning and abstract problem solving, which are less studied and existing studies only focus on simple graphics. To evaluate the performance of VLMs in complex graphic reasoning, we propose ReasonBench, the first evaluation benchmark focused on structured graphic reasoning tasks, which includes 1,613 questions from real-world intelligence tests. ReasonBench covers reasoning dimensions related to location, attribute, quantity, and multi-element tasks, providing a comprehensive evaluation of the performance of VLMs in spatial, relational, and abstract reasoning capabilities. We benchmark 11 mainstream VLMs (including closed-source and open-source models) and reveal significant limitations of current models. Based on these findings, we propose a dual optimization strategy: Diagrammatic Reasoning Chain (DiaCoT) enhances the interpretability of reasoning by decomposing layers, and ReasonTune enhances the task adaptability of model reasoning through training, all of which improves VLM performance by 33.5\%. All experimental data and code are in the repository: https://huggingface.co/datasets/cistine/ReasonBench.
title Oedipus and the Sphinx: Benchmarking and Improving Visual Language Models for Complex Graphic Reasoning
topic Artificial Intelligence
url https://arxiv.org/abs/2508.00323