GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xia, Renqiu, Li, Mingsheng, Ye, Hancheng, Wu, Wenjie, Zhou, Hongbin, Yuan, Jiakang, Peng, Tianshuo, Cai, Xinyu, Yan, Xiangchao, Wang, Bin, He, Conghui, Shi, Botian, Chen, Tao, Yan, Junchi, Zhang, Bo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915096788205568
author Xia, Renqiu
Li, Mingsheng
Ye, Hancheng
Wu, Wenjie
Zhou, Hongbin
Yuan, Jiakang
Peng, Tianshuo
Cai, Xinyu
Yan, Xiangchao
Wang, Bin
He, Conghui
Shi, Botian
Chen, Tao
Yan, Junchi
Zhang, Bo
author_facet Xia, Renqiu
Li, Mingsheng
Ye, Hancheng
Wu, Wenjie
Zhou, Hongbin
Yuan, Jiakang
Peng, Tianshuo
Cai, Xinyu
Yan, Xiangchao
Wang, Bin
He, Conghui
Shi, Botian
Chen, Tao
Yan, Junchi
Zhang, Bo
contents Despite their proficiency in general tasks, Multi-modal Large Language Models (MLLMs) struggle with automatic Geometry Problem Solving (GPS), which demands understanding diagrams, interpreting symbols, and performing complex reasoning. This limitation arises from their pre-training on natural images and texts, along with the lack of automated verification in the problem-solving process. Besides, current geometric specialists are limited by their task-specific designs, making them less effective for broader geometric problems. To this end, we present GeoX, a multi-modal large model focusing on geometric understanding and reasoning tasks. Given the significant differences between geometric diagram-symbol and natural image-text, we introduce unimodal pre-training to develop a diagram encoder and symbol decoder, enhancing the understanding of geometric images and corpora. Furthermore, we introduce geometry-language alignment, an effective pre-training paradigm that bridges the modality gap between unimodal geometric experts. We propose a Generator-And-Sampler Transformer (GS-Former) to generate discriminative queries and eliminate uninformative representations from unevenly distributed geometric signals. Finally, GeoX benefits from visual instruction tuning, empowering it to take geometric images and questions as input and generate verifiable solutions. Experiments show that GeoX outperforms both generalists and geometric specialists on publicly recognized benchmarks, such as GeoQA, UniGeo, Geometry3K, and PGPS9k.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11863
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
Xia, Renqiu
Li, Mingsheng
Ye, Hancheng
Wu, Wenjie
Zhou, Hongbin
Yuan, Jiakang
Peng, Tianshuo
Cai, Xinyu
Yan, Xiangchao
Wang, Bin
He, Conghui
Shi, Botian
Chen, Tao
Yan, Junchi
Zhang, Bo
Computer Vision and Pattern Recognition
Computation and Language
Despite their proficiency in general tasks, Multi-modal Large Language Models (MLLMs) struggle with automatic Geometry Problem Solving (GPS), which demands understanding diagrams, interpreting symbols, and performing complex reasoning. This limitation arises from their pre-training on natural images and texts, along with the lack of automated verification in the problem-solving process. Besides, current geometric specialists are limited by their task-specific designs, making them less effective for broader geometric problems. To this end, we present GeoX, a multi-modal large model focusing on geometric understanding and reasoning tasks. Given the significant differences between geometric diagram-symbol and natural image-text, we introduce unimodal pre-training to develop a diagram encoder and symbol decoder, enhancing the understanding of geometric images and corpora. Furthermore, we introduce geometry-language alignment, an effective pre-training paradigm that bridges the modality gap between unimodal geometric experts. We propose a Generator-And-Sampler Transformer (GS-Former) to generate discriminative queries and eliminate uninformative representations from unevenly distributed geometric signals. Finally, GeoX benefits from visual instruction tuning, empowering it to take geometric images and questions as input and generate verifiable solutions. Experiments show that GeoX outperforms both generalists and geometric specialists on publicly recognized benchmarks, such as GeoQA, UniGeo, Geometry3K, and PGPS9k.
title GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2412.11863