Saved in:
Bibliographic Details
Main Authors: Jia, Mengzhao, Zhang, Zhihan, Yu, Wenhao, Jiao, Fangkai, Jiang, Meng
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2404.14604
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916223780913152
author Jia, Mengzhao
Zhang, Zhihan
Yu, Wenhao
Jiao, Fangkai
Jiang, Meng
author_facet Jia, Mengzhao
Zhang, Zhihan
Yu, Wenhao
Jiao, Fangkai
Jiang, Meng
contents Open-source multimodal large language models (MLLMs) excel in various tasks involving textual and visual inputs but still struggle with complex multimodal mathematical reasoning, lagging behind proprietary models like GPT-4V(ision) and Gemini-Pro. Although fine-tuning with intermediate steps (i.e., rationales) elicits some mathematical reasoning skills, the resulting models still fall short in visual comprehension due to inadequate visual-centric supervision, which leads to inaccurate interpretation of math figures. To address this issue, we propose a two-step training pipeline VCAR, which emphasizes the Visual Comprehension training in Addition to mathematical Reasoning learning. It first improves the visual comprehension ability of MLLMs through the visual description generation task, followed by another training step on generating rationales with the assistance of descriptions. Experimental results on two popular benchmarks demonstrate that VCAR substantially outperforms baseline methods solely relying on rationale supervision, especially on problems with high visual demands.
format Preprint
id arxiv_https___arxiv_org_abs_2404_14604
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Describe-then-Reason: Improving Multimodal Mathematical Reasoning through Visual Comprehension Training
Jia, Mengzhao
Zhang, Zhihan
Yu, Wenhao
Jiao, Fangkai
Jiang, Meng
Computation and Language
Open-source multimodal large language models (MLLMs) excel in various tasks involving textual and visual inputs but still struggle with complex multimodal mathematical reasoning, lagging behind proprietary models like GPT-4V(ision) and Gemini-Pro. Although fine-tuning with intermediate steps (i.e., rationales) elicits some mathematical reasoning skills, the resulting models still fall short in visual comprehension due to inadequate visual-centric supervision, which leads to inaccurate interpretation of math figures. To address this issue, we propose a two-step training pipeline VCAR, which emphasizes the Visual Comprehension training in Addition to mathematical Reasoning learning. It first improves the visual comprehension ability of MLLMs through the visual description generation task, followed by another training step on generating rationales with the assistance of descriptions. Experimental results on two popular benchmarks demonstrate that VCAR substantially outperforms baseline methods solely relying on rationale supervision, especially on problems with high visual demands.
title Describe-then-Reason: Improving Multimodal Mathematical Reasoning through Visual Comprehension Training
topic Computation and Language
url https://arxiv.org/abs/2404.14604