Improving Multimodal LLMs Ability In Geometry Problem Solving, Reasoning, And Multistep Scoring

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Anand, Avinash, Jaiswal, Raj, Dharmadhikari, Abhishek, Marathe, Atharva, Popat, Harsh Parimal, Mital, Harshil, Prasad, Kritarth, Shah, Rajiv Ratn, Zimmermann, Roger
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915042516008960
author Anand, Avinash
Jaiswal, Raj
Dharmadhikari, Abhishek
Marathe, Atharva
Popat, Harsh Parimal
Mital, Harshil
Prasad, Kritarth
Shah, Rajiv Ratn
Zimmermann, Roger
author_facet Anand, Avinash
Jaiswal, Raj
Dharmadhikari, Abhishek
Marathe, Atharva
Popat, Harsh Parimal
Mital, Harshil
Prasad, Kritarth
Shah, Rajiv Ratn
Zimmermann, Roger
contents This paper presents GPSM4K, a comprehensive geometry multimodal dataset tailored to augment the problem-solving capabilities of Large Vision Language Models (LVLMs). GPSM4K encompasses 2157 multimodal question-answer pairs manually extracted from mathematics textbooks spanning grades 7-12 and is further augmented to 5340 problems, consisting of both numerical and theorem-proving questions. In contrast to PGPS9k, Geometry3K, and Geo170K which feature only objective-type questions, GPSM4K offers detailed step-by-step solutions in a consistent format, facilitating a comprehensive evaluation of problem-solving approaches. This dataset serves as an excellent benchmark for assessing the geometric reasoning capabilities of LVLMs. Evaluation of our test set shows that there is scope for improvement needed in open-source language models in geometry problem-solving. Finetuning on our training set increases the geometry problem-solving capabilities of models. Further, We also evaluate the effectiveness of techniques such as image captioning and Retrieval Augmentation generation (RAG) on model performance. We leveraged LLM to automate the task of final answer evaluation by providing ground truth and predicted solutions. This research will help to assess and improve the geometric reasoning capabilities of LVLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00846
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Multimodal LLMs Ability In Geometry Problem Solving, Reasoning, And Multistep Scoring
Anand, Avinash
Jaiswal, Raj
Dharmadhikari, Abhishek
Marathe, Atharva
Popat, Harsh Parimal
Mital, Harshil
Prasad, Kritarth
Shah, Rajiv Ratn
Zimmermann, Roger
Artificial Intelligence
This paper presents GPSM4K, a comprehensive geometry multimodal dataset tailored to augment the problem-solving capabilities of Large Vision Language Models (LVLMs). GPSM4K encompasses 2157 multimodal question-answer pairs manually extracted from mathematics textbooks spanning grades 7-12 and is further augmented to 5340 problems, consisting of both numerical and theorem-proving questions. In contrast to PGPS9k, Geometry3K, and Geo170K which feature only objective-type questions, GPSM4K offers detailed step-by-step solutions in a consistent format, facilitating a comprehensive evaluation of problem-solving approaches. This dataset serves as an excellent benchmark for assessing the geometric reasoning capabilities of LVLMs. Evaluation of our test set shows that there is scope for improvement needed in open-source language models in geometry problem-solving. Finetuning on our training set increases the geometry problem-solving capabilities of models. Further, We also evaluate the effectiveness of techniques such as image captioning and Retrieval Augmentation generation (RAG) on model performance. We leveraged LLM to automate the task of final answer evaluation by providing ground truth and predicted solutions. This research will help to assess and improve the geometric reasoning capabilities of LVLMs.
title Improving Multimodal LLMs Ability In Geometry Problem Solving, Reasoning, And Multistep Scoring
topic Artificial Intelligence
url https://arxiv.org/abs/2412.00846