ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Lihong, Li, Liangqi, Feng, Weiwei, Wu, Jiamin, Miao, Changtao, Wu, Tieru, Ma, Rui, Zhang, Bo, Li, Zhe
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918372063576064
author Wang, Lihong
Li, Liangqi
Feng, Weiwei
Wu, Jiamin
Miao, Changtao
Wu, Tieru
Ma, Rui
Zhang, Bo
Li, Zhe
author_facet Wang, Lihong
Li, Liangqi
Feng, Weiwei
Wu, Jiamin
Miao, Changtao
Wu, Tieru
Ma, Rui
Zhang, Bo
Li, Zhe
contents CoT has significantly enhanced the reasoning ability of LLMs while it faces challenges when extended to multimodal domains, particularly in mathematical tasks. Existing MLLMs typically perform textual reasoning solely from a single static mathematical image, overlooking dynamic visual acquisition during reasoning. In contrast, humans repeatedly examine visual image and employ step-by-step reasoning to prove intermediate propositions. This strategy of decomposing the problem-solving process into key logical nodes adheres to Miller's Law in cognitive science. Inspired by this insight, we propose a ViRC framework for multimodal mathematical tasks, introducing a Reason Chunking mechanism that structures multimodal mathematical CoT into consecutive Critical Reasoning Units (CRUs) to simulate human expert problem-solving patterns. CRUs ensure intra-unit textual coherence for intermediate proposition verification while integrating visual information across units to generate subsequent propositions and support structured reasoning. To this end, we present CRUX dataset by using three visual tools and four reasoning patterns to provide explicitly annotated CRUs across multiple reasoning paths for each mathematical problem. Leveraging the CRUX dataset, we propose a progressive training strategy inspired by human cognitive learning, which includes Instructional SFT, Practice SFT, and Strategic RL, aimed at further strengthening the Reason Chunking ability of the model. The resulting ViRC-7B model achieves a 18.8% average improvement over baselines across multiple mathematical benchmarks. Code is available at https://github.com/Leon-LihongWang/ViRC.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14654
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
Wang, Lihong
Li, Liangqi
Feng, Weiwei
Wu, Jiamin
Miao, Changtao
Wu, Tieru
Ma, Rui
Zhang, Bo
Li, Zhe
Computer Vision and Pattern Recognition
CoT has significantly enhanced the reasoning ability of LLMs while it faces challenges when extended to multimodal domains, particularly in mathematical tasks. Existing MLLMs typically perform textual reasoning solely from a single static mathematical image, overlooking dynamic visual acquisition during reasoning. In contrast, humans repeatedly examine visual image and employ step-by-step reasoning to prove intermediate propositions. This strategy of decomposing the problem-solving process into key logical nodes adheres to Miller's Law in cognitive science. Inspired by this insight, we propose a ViRC framework for multimodal mathematical tasks, introducing a Reason Chunking mechanism that structures multimodal mathematical CoT into consecutive Critical Reasoning Units (CRUs) to simulate human expert problem-solving patterns. CRUs ensure intra-unit textual coherence for intermediate proposition verification while integrating visual information across units to generate subsequent propositions and support structured reasoning. To this end, we present CRUX dataset by using three visual tools and four reasoning patterns to provide explicitly annotated CRUs across multiple reasoning paths for each mathematical problem. Leveraging the CRUX dataset, we propose a progressive training strategy inspired by human cognitive learning, which includes Instructional SFT, Practice SFT, and Strategic RL, aimed at further strengthening the Reason Chunking ability of the model. The resulting ViRC-7B model achieves a 18.8% average improvement over baselines across multiple mathematical benchmarks. Code is available at https://github.com/Leon-LihongWang/ViRC.
title ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.14654