GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Feng, Yuan, Yang, Yue, He, Xiaohan, Zhao, Jiatong, Chen, Jianlong, Chen, Zijun, Fu, Daocheng, Liu, Qi, Xia, Renqiu, Zhang, Bo, Yan, Junchi
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917177285672960
author Feng, Yuan
Yang, Yue
He, Xiaohan
Zhao, Jiatong
Chen, Jianlong
Chen, Zijun
Fu, Daocheng
Liu, Qi
Xia, Renqiu
Zhang, Bo
Yan, Junchi
author_facet Feng, Yuan
Yang, Yue
He, Xiaohan
Zhao, Jiatong
Chen, Jianlong
Chen, Zijun
Fu, Daocheng
Liu, Qi
Xia, Renqiu
Zhang, Bo
Yan, Junchi
contents Geometric problem solving constitutes a critical branch of mathematical reasoning, requiring precise analysis of shapes and spatial relationships. Current evaluations of geometric reasoning in vision-language models (VLMs) face limitations, including the risk of test data contamination from textbook-based benchmarks, overemphasis on final answers over reasoning processes, and insufficient diagnostic granularity. To address these issues, we present GeoBench, a hierarchical benchmark featuring four reasoning levels in geometric problem-solving: Visual Perception, Goal-Oriented Planning, Rigorous Theorem Application, and Self-Reflective Backtracking. Through six formally verified tasks generated via TrustGeoGen, we systematically assess capabilities ranging from attribute extraction to logical error correction. Experiments reveal that while reasoning models like OpenAI-o3 outperform general MLLMs, performance declines significantly with increasing task complexity. Key findings demonstrate that sub-goal decomposition and irrelevant premise filtering critically influence final problem-solving accuracy, whereas Chain-of-Thought prompting unexpectedly degrades performance in some tasks. These findings establish GeoBench as a comprehensive benchmark while offering actionable guidelines for developing geometric problem-solving systems.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24119
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation
Feng, Yuan
Yang, Yue
He, Xiaohan
Zhao, Jiatong
Chen, Jianlong
Chen, Zijun
Fu, Daocheng
Liu, Qi
Xia, Renqiu
Zhang, Bo
Yan, Junchi
Computer Vision and Pattern Recognition
Geometric problem solving constitutes a critical branch of mathematical reasoning, requiring precise analysis of shapes and spatial relationships. Current evaluations of geometric reasoning in vision-language models (VLMs) face limitations, including the risk of test data contamination from textbook-based benchmarks, overemphasis on final answers over reasoning processes, and insufficient diagnostic granularity. To address these issues, we present GeoBench, a hierarchical benchmark featuring four reasoning levels in geometric problem-solving: Visual Perception, Goal-Oriented Planning, Rigorous Theorem Application, and Self-Reflective Backtracking. Through six formally verified tasks generated via TrustGeoGen, we systematically assess capabilities ranging from attribute extraction to logical error correction. Experiments reveal that while reasoning models like OpenAI-o3 outperform general MLLMs, performance declines significantly with increasing task complexity. Key findings demonstrate that sub-goal decomposition and irrelevant premise filtering critically influence final problem-solving accuracy, whereas Chain-of-Thought prompting unexpectedly degrades performance in some tasks. These findings establish GeoBench as a comprehensive benchmark while offering actionable guidelines for developing geometric problem-solving systems.
title GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.24119