Saved in:
Bibliographic Details
Main Authors: Xi, Zhiheng, Li, Guanyu, Fan, Yutao, Guo, Honglin, Liu, Yufang, Fan, Xiaoran, Liu, Jiaqi, Ding, Jingchao, Zuo, Wangmeng, Yin, Zhenfei, Bai, Lei, Ji, Tao, Gui, Tao, Zhang, Qi, Torr, Philip, Huang, Xuanjing
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.03483
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908439833214976
author Xi, Zhiheng
Li, Guanyu
Fan, Yutao
Guo, Honglin
Liu, Yufang
Fan, Xiaoran
Liu, Jiaqi
Ding, Jingchao
Zuo, Wangmeng
Yin, Zhenfei
Bai, Lei
Ji, Tao
Gui, Tao
Zhang, Qi
Torr, Philip
Huang, Xuanjing
author_facet Xi, Zhiheng
Li, Guanyu
Fan, Yutao
Guo, Honglin
Liu, Yufang
Fan, Xiaoran
Liu, Jiaqi
Ding, Jingchao
Zuo, Wangmeng
Yin, Zhenfei
Bai, Lei
Ji, Tao
Gui, Tao
Zhang, Qi
Torr, Philip
Huang, Xuanjing
contents In this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 110k college-level questions spanning 300 UNESCO-defined subjects, spanning diverse formats-multiple-choice, fill-in-the-blank, and open-ended QA-and sourced from both print and digital media such as books, exams, and quizzes. All data are curated and filtered via a human-in-the-loop and scalable framework, and each instance is paired with a high-quality reasoning path. The dataset is organized into two parts: BMMR-Eval that comprises 20,458 high-quality instances to comprehensively assess LMMs' knowledge and reasoning across multiple disciplines in both Chinese and English; and BMMR-Train that contains 88,991 instances to support further research and development, extending the current focus on mathematical reasoning to diverse disciplines and domains. In addition, we propose the process-based multi-discipline verifier (i.e., BMMR-Verifier) for accurate and fine-grained evaluation of reasoning paths. Extensive experiments on 24 models reveal that (i) even SOTA models (e.g., o3 and Gemini-2.5-Pro) leave substantial headroom on BMMR-Eval; (ii) reasoning models exhibit discipline bias and outperform LMMs only on specific subjects; (iii) open-source models still trail their proprietary counterparts; and (iv) fine-tuning on BMMR-Train narrows this gap. Additionally, we conduct reasoning-chain analyses using BMMR-Verifier and other in-depth studies, uncovering the challenges LMMs currently face in multidisciplinary reasoning. We will release the data, and we hope our work can offer insights and contributions to the community.
format Preprint
id arxiv_https___arxiv_org_abs_2507_03483
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
Xi, Zhiheng
Li, Guanyu
Fan, Yutao
Guo, Honglin
Liu, Yufang
Fan, Xiaoran
Liu, Jiaqi
Ding, Jingchao
Zuo, Wangmeng
Yin, Zhenfei
Bai, Lei
Ji, Tao
Gui, Tao
Zhang, Qi
Torr, Philip
Huang, Xuanjing
Computation and Language
Artificial Intelligence
In this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 110k college-level questions spanning 300 UNESCO-defined subjects, spanning diverse formats-multiple-choice, fill-in-the-blank, and open-ended QA-and sourced from both print and digital media such as books, exams, and quizzes. All data are curated and filtered via a human-in-the-loop and scalable framework, and each instance is paired with a high-quality reasoning path. The dataset is organized into two parts: BMMR-Eval that comprises 20,458 high-quality instances to comprehensively assess LMMs' knowledge and reasoning across multiple disciplines in both Chinese and English; and BMMR-Train that contains 88,991 instances to support further research and development, extending the current focus on mathematical reasoning to diverse disciplines and domains. In addition, we propose the process-based multi-discipline verifier (i.e., BMMR-Verifier) for accurate and fine-grained evaluation of reasoning paths. Extensive experiments on 24 models reveal that (i) even SOTA models (e.g., o3 and Gemini-2.5-Pro) leave substantial headroom on BMMR-Eval; (ii) reasoning models exhibit discipline bias and outperform LMMs only on specific subjects; (iii) open-source models still trail their proprietary counterparts; and (iv) fine-tuning on BMMR-Train narrows this gap. Additionally, we conduct reasoning-chain analyses using BMMR-Verifier and other in-depth studies, uncovering the challenges LMMs currently face in multidisciplinary reasoning. We will release the data, and we hope our work can offer insights and contributions to the community.
title BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.03483