A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yan, Yibo, Su, Jiamin, He, Jianxiang, Fu, Fangteng, Zheng, Xu, Lyu, Yuanhuiyi, Wang, Kun, Wang, Shen, Wen, Qingsong, Hu, Xuming
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915294407032832
author Yan, Yibo
Su, Jiamin
He, Jianxiang
Fu, Fangteng
Zheng, Xu
Lyu, Yuanhuiyi
Wang, Kun
Wang, Shen
Wen, Qingsong
Hu, Xuming
author_facet Yan, Yibo
Su, Jiamin
He, Jianxiang
Fu, Fangteng
Zheng, Xu
Lyu, Yuanhuiyi
Wang, Kun
Wang, Shen
Wen, Qingsong
Hu, Xuming
contents Mathematical reasoning, a core aspect of human cognition, is vital across many domains, from educational problem-solving to scientific advancements. As artificial general intelligence (AGI) progresses, integrating large language models (LLMs) with mathematical reasoning tasks is becoming increasingly significant. This survey provides the first comprehensive analysis of mathematical reasoning in the era of multimodal large language models (MLLMs). We review over 200 studies published since 2021, and examine the state-of-the-art developments in Math-LLMs, with a focus on multimodal settings. We categorize the field into three dimensions: benchmarks, methodologies, and challenges. In particular, we explore multimodal mathematical reasoning pipeline, as well as the role of (M)LLMs and the associated methodologies. Finally, we identify five major challenges hindering the realization of AGI in this domain, offering insights into the future direction for enhancing multimodal reasoning capabilities. This survey serves as a critical resource for the research community in advancing the capabilities of LLMs to tackle complex multimodal reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11936
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges
Yan, Yibo
Su, Jiamin
He, Jianxiang
Fu, Fangteng
Zheng, Xu
Lyu, Yuanhuiyi
Wang, Kun
Wang, Shen
Wen, Qingsong
Hu, Xuming
Computation and Language
Mathematical reasoning, a core aspect of human cognition, is vital across many domains, from educational problem-solving to scientific advancements. As artificial general intelligence (AGI) progresses, integrating large language models (LLMs) with mathematical reasoning tasks is becoming increasingly significant. This survey provides the first comprehensive analysis of mathematical reasoning in the era of multimodal large language models (MLLMs). We review over 200 studies published since 2021, and examine the state-of-the-art developments in Math-LLMs, with a focus on multimodal settings. We categorize the field into three dimensions: benchmarks, methodologies, and challenges. In particular, we explore multimodal mathematical reasoning pipeline, as well as the role of (M)LLMs and the associated methodologies. Finally, we identify five major challenges hindering the realization of AGI in this domain, offering insights into the future direction for enhancing multimodal reasoning capabilities. This survey serves as a critical resource for the research community in advancing the capabilities of LLMs to tackle complex multimodal reasoning tasks.
title A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges
topic Computation and Language
url https://arxiv.org/abs/2412.11936