MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Renrui, Wei, Xinyu, Jiang, Dongzhi, Guo, Ziyu, Li, Shicheng, Zhang, Yichi, Tong, Chengzhuo, Liu, Jiaming, Zhou, Aojun, Wei, Bin, Zhang, Shanghang, Gao, Peng, Li, Chunyuan, Li, Hongsheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915002295779328
author Zhang, Renrui
Wei, Xinyu
Jiang, Dongzhi
Guo, Ziyu
Li, Shicheng
Zhang, Yichi
Tong, Chengzhuo
Liu, Jiaming
Zhou, Aojun
Wei, Bin
Zhang, Shanghang
Gao, Peng
Li, Chunyuan
Li, Hongsheng
author_facet Zhang, Renrui
Wei, Xinyu
Jiang, Dongzhi
Guo, Ziyu
Li, Shicheng
Zhang, Yichi
Tong, Chengzhuo
Liu, Jiaming
Zhou, Aojun
Wei, Bin
Zhang, Shanghang
Gao, Peng
Li, Chunyuan
Li, Hongsheng
contents The mathematical capabilities of Multi-modal Large Language Models (MLLMs) remain under-explored with three areas to be improved: visual encoding of math diagrams, diagram-language alignment, and chain-of-thought (CoT) reasoning. This draws forth an urgent demand for an effective training paradigm and a large-scale, comprehensive dataset with detailed CoT rationales, which is challenging to collect and costly to annotate manually. To tackle this issue, we propose MAVIS, a MAthematical VISual instruction tuning pipeline for MLLMs, featuring an automatic data engine to efficiently create mathematical visual datasets. We design the data generation process to be entirely independent of human intervention or GPT API usage, while ensuring the diagram-caption correspondence, question-answer correctness, and CoT reasoning quality. With this approach, we curate two datasets, MAVIS-Caption (558K diagram-caption pairs) and MAVIS-Instruct (834K visual math problems with CoT rationales), and propose four progressive stages for training MLLMs from scratch. First, we utilize MAVIS-Caption to fine-tune a math-specific vision encoder (CLIP-Math) through contrastive learning, tailored for improved diagram visual encoding. Second, we also leverage MAVIS-Caption to align the CLIP-Math with a large language model (LLM) by a projection layer, enhancing vision-language alignment in mathematical domains. Third, we adopt MAVIS-Instruct to perform the instruction tuning for robust problem-solving skills, and term the resulting model as MAVIS-7B. Fourth, we apply Direct Preference Optimization (DPO) to enhance the CoT capabilities of our model, further refining its step-wise reasoning performance. Code and data will be released at https://github.com/ZrrSkywalker/MAVIS
format Preprint
id arxiv_https___arxiv_org_abs_2407_08739
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
Zhang, Renrui
Wei, Xinyu
Jiang, Dongzhi
Guo, Ziyu
Li, Shicheng
Zhang, Yichi
Tong, Chengzhuo
Liu, Jiaming
Zhou, Aojun
Wei, Bin
Zhang, Shanghang
Gao, Peng
Li, Chunyuan
Li, Hongsheng
Computer Vision and Pattern Recognition
The mathematical capabilities of Multi-modal Large Language Models (MLLMs) remain under-explored with three areas to be improved: visual encoding of math diagrams, diagram-language alignment, and chain-of-thought (CoT) reasoning. This draws forth an urgent demand for an effective training paradigm and a large-scale, comprehensive dataset with detailed CoT rationales, which is challenging to collect and costly to annotate manually. To tackle this issue, we propose MAVIS, a MAthematical VISual instruction tuning pipeline for MLLMs, featuring an automatic data engine to efficiently create mathematical visual datasets. We design the data generation process to be entirely independent of human intervention or GPT API usage, while ensuring the diagram-caption correspondence, question-answer correctness, and CoT reasoning quality. With this approach, we curate two datasets, MAVIS-Caption (558K diagram-caption pairs) and MAVIS-Instruct (834K visual math problems with CoT rationales), and propose four progressive stages for training MLLMs from scratch. First, we utilize MAVIS-Caption to fine-tune a math-specific vision encoder (CLIP-Math) through contrastive learning, tailored for improved diagram visual encoding. Second, we also leverage MAVIS-Caption to align the CLIP-Math with a large language model (LLM) by a projection layer, enhancing vision-language alignment in mathematical domains. Third, we adopt MAVIS-Instruct to perform the instruction tuning for robust problem-solving skills, and term the resulting model as MAVIS-7B. Fourth, we apply Direct Preference Optimization (DPO) to enhance the CoT capabilities of our model, further refining its step-wise reasoning performance. Code and data will be released at https://github.com/ZrrSkywalker/MAVIS
title MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.08739