GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Yuan, Yang, Yue, He, Xiaohan, Zhao, Jiatong, Chen, Jianlong, Chen, Zijun, Fu, Daocheng, Liu, Qi, Xia, Renqiu, Zhang, Bo, Yan, Junchi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving
by: Fu, Daocheng, et al.
Published: (2025)
by: Fu, Daocheng, et al.
Published: (2025)
Milestones over Outcome: Unlocking Geometric Reasoning with Sub-Goal Verifiable Reward
by: Chen, Jianlong, et al.
Published: (2026)
by: Chen, Jianlong, et al.
Published: (2026)
Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving
by: Liu, Qi, et al.
Published: (2025)
by: Liu, Qi, et al.
Published: (2025)
GeoBench: Benchmarking and Analyzing Monocular Geometry Estimation Models
by: Ge, Yongtao, et al.
Published: (2024)
by: Ge, Yongtao, et al.
Published: (2024)
GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
by: Xia, Renqiu, et al.
Published: (2024)
by: Xia, Renqiu, et al.
Published: (2024)
On the Evaluation and Refinement of Vision-Language Instruction Tuning Datasets
by: Liao, Ning, et al.
Published: (2023)
by: Liao, Ning, et al.
Published: (2023)
Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy
by: Ye, Hancheng, et al.
Published: (2024)
by: Ye, Hancheng, et al.
Published: (2024)
ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
by: Xia, Renqiu, et al.
Published: (2024)
by: Xia, Renqiu, et al.
Published: (2024)
SurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing
by: Yan, Xiangchao, et al.
Published: (2025)
by: Yan, Xiangchao, et al.
Published: (2025)
On the Emergence of Cross-Task Linearity in the Pretraining-Finetuning Paradigm
by: Zhou, Zhanpeng, et al.
Published: (2024)
by: Zhou, Zhanpeng, et al.
Published: (2024)
PuzzleBench: A Fully Dynamic Evaluation Framework for Large Multimodal Models on Puzzle Solving
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
FormalGeo: An Extensible Formalized Framework for Olympiad Geometric Problem Solving
by: Zhang, Xiaokai, et al.
Published: (2023)
by: Zhang, Xiaokai, et al.
Published: (2023)
Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural Integration
by: Pan, Yicheng, et al.
Published: (2025)
by: Pan, Yicheng, et al.
Published: (2025)
Rethinking the Potential of Multimodality in Collaborative Problem Solving Diagnosis with Large Language Models
by: Wong, K., et al.
Published: (2025)
by: Wong, K., et al.
Published: (2025)
Concise Geometric Description as a Bridge: Unleashing the Potential of LLM for Plane Geometry Problem Solving
by: Wang, Jingyun, et al.
Published: (2026)
by: Wang, Jingyun, et al.
Published: (2026)
Hierarchical Attention Generates Better Proofs
by: Chen, Jianlong, et al.
Published: (2025)
by: Chen, Jianlong, et al.
Published: (2025)
SPD-Faith Bench: Diagnosing and Improving Faithfulness in Chain-of-Thought for Multimodal Large Language Models
by: Lv, Weijiang, et al.
Published: (2026)
by: Lv, Weijiang, et al.
Published: (2026)
GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning
by: Xu, Liangyu, et al.
Published: (2025)
by: Xu, Liangyu, et al.
Published: (2025)
Hilbert-Geo: Solving Solid Geometric Problems by Neural-Symbolic Reasoning
by: Xu, Ruoran, et al.
Published: (2026)
by: Xu, Ruoran, et al.
Published: (2026)
SO-Bench: A Structural Output Evaluation of Multimodal LLMs
by: Feng, Di, et al.
Published: (2025)
by: Feng, Di, et al.
Published: (2025)
MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models
by: Jin, Bohan, et al.
Published: (2025)
by: Jin, Bohan, et al.
Published: (2025)
SE-Merging: A Self-Enhanced Approach for Dynamic Model Merging
by: Chen, Zijun, et al.
Published: (2025)
by: Chen, Zijun, et al.
Published: (2025)
DriveVGGT: Calibration-Constrained Visual Geometry Transformers for Multi-Camera Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2025)
by: Jia, Xiaosong, et al.
Published: (2025)
Efficient, Adaptive Near-Field Beam Training based on Linear Bandit
by: Liu, Junchi, et al.
Published: (2026)
by: Liu, Junchi, et al.
Published: (2026)
Algorithm Research of ELMo Word Embedding and Deep Learning Multimodal Transformer in Image Description
by: Cheng, Xiaohan, et al.
Published: (2024)
by: Cheng, Xiaohan, et al.
Published: (2024)
Can LLM Assist in the Evaluation of the Quality of Machine Learning Explanations?
by: Wang, Bo, et al.
Published: (2025)
by: Wang, Bo, et al.
Published: (2025)
StructChart: On the Schema, Metric, and Augmentation for Visual Chart Understanding
by: Xia, Renqiu, et al.
Published: (2023)
by: Xia, Renqiu, et al.
Published: (2023)
EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving
by: Zhou, Xiyuan, et al.
Published: (2025)
by: Zhou, Xiyuan, et al.
Published: (2025)
ReSimAD: Zero-Shot 3D Domain Transfer for Autonomous Driving with Source Reconstruction and Target Simulation
by: Zhang, Bo, et al.
Published: (2023)
by: Zhang, Bo, et al.
Published: (2023)
SPOT: Scalable 3D Pre-training via Occupancy Prediction for Learning Transferable 3D Representations
by: Yan, Xiangchao, et al.
Published: (2023)
by: Yan, Xiangchao, et al.
Published: (2023)
Multipole phases in a type of spin/fermion ladders with local conserved quantities and generalizations
by: Fu, Jianlong
Published: (2025)
by: Fu, Jianlong
Published: (2025)
QuantumBench: A Benchmark for Quantum Problem Solving
by: Minami, Shunya, et al.
Published: (2025)
by: Minami, Shunya, et al.
Published: (2025)
GeoR-Bench: Evaluating Geoscience Visual Reasoning
by: Zheng, Yushuo, et al.
Published: (2026)
by: Zheng, Yushuo, et al.
Published: (2026)
CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction
by: Ma, Yinghao, et al.
Published: (2026)
by: Ma, Yinghao, et al.
Published: (2026)
MuteBench: Modality Unavailability Tolerance Evaluation for Incomplete Multimodal Fusion
by: Zheng, Wugeng, et al.
Published: (2026)
by: Zheng, Wugeng, et al.
Published: (2026)
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
by: Deng, Andong, et al.
Published: (2025)
by: Deng, Andong, et al.
Published: (2025)
MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks
by: Chen, Jiacheng, et al.
Published: (2024)
by: Chen, Jiacheng, et al.
Published: (2024)
GeoGramBench: Benchmarking the Geometric Program Reasoning in Modern LLMs
by: Luo, Shixian, et al.
Published: (2025)
by: Luo, Shixian, et al.
Published: (2025)
GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models
by: Xie, Qinghongbing, et al.
Published: (2025)
by: Xie, Qinghongbing, et al.
Published: (2025)
Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching
by: Wang, Bin, et al.
Published: (2024)
by: Wang, Bin, et al.
Published: (2024)
Similar Items
-
TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving
by: Fu, Daocheng, et al.
Published: (2025) -
Milestones over Outcome: Unlocking Geometric Reasoning with Sub-Goal Verifiable Reward
by: Chen, Jianlong, et al.
Published: (2026) -
Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving
by: Liu, Qi, et al.
Published: (2025) -
GeoBench: Benchmarking and Analyzing Monocular Geometry Estimation Models
by: Ge, Yongtao, et al.
Published: (2024) -
GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
by: Xia, Renqiu, et al.
Published: (2024)