ClozeMath: Improving Mathematical Reasoning in Language Models by Learning to Fill Equations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pham, Quang Hieu, Nguyen, Thuy Duong, Pham, Tung, Luu, Anh Tuan, Nguyen, Dat Quoc
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918045441589248
author Pham, Quang Hieu
Nguyen, Thuy Duong
Pham, Tung
Luu, Anh Tuan
Nguyen, Dat Quoc
author_facet Pham, Quang Hieu
Nguyen, Thuy Duong
Pham, Tung
Luu, Anh Tuan
Nguyen, Dat Quoc
contents The capabilities of large language models (LLMs) have been enhanced by training on data that reflects human thought processes, such as the Chain-of-Thought format. However, evidence suggests that the conventional scheme of next-word prediction may not fully capture how humans learn to think. Inspired by how humans generalize mathematical reasoning, we propose a new approach named ClozeMath to fine-tune LLMs for mathematical reasoning. Our ClozeMath involves a text-infilling task that predicts masked equations from a given solution, analogous to cloze exercises used in human learning. Experiments on GSM8K, MATH, and GSM-Symbolic show that ClozeMath surpasses the strong baseline Masked Thought in performance and robustness, with two test-time scaling decoding algorithms, Beam Search and Chain-of-Thought decoding. Additionally, we conduct an ablation study to analyze the effects of various architectural and implementation choices on our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03763
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ClozeMath: Improving Mathematical Reasoning in Language Models by Learning to Fill Equations
Pham, Quang Hieu
Nguyen, Thuy Duong
Pham, Tung
Luu, Anh Tuan
Nguyen, Dat Quoc
Computation and Language
The capabilities of large language models (LLMs) have been enhanced by training on data that reflects human thought processes, such as the Chain-of-Thought format. However, evidence suggests that the conventional scheme of next-word prediction may not fully capture how humans learn to think. Inspired by how humans generalize mathematical reasoning, we propose a new approach named ClozeMath to fine-tune LLMs for mathematical reasoning. Our ClozeMath involves a text-infilling task that predicts masked equations from a given solution, analogous to cloze exercises used in human learning. Experiments on GSM8K, MATH, and GSM-Symbolic show that ClozeMath surpasses the strong baseline Masked Thought in performance and robustness, with two test-time scaling decoding algorithms, Beam Search and Chain-of-Thought decoding. Additionally, we conduct an ablation study to analyze the effects of various architectural and implementation choices on our approach.
title ClozeMath: Improving Mathematical Reasoning in Language Models by Learning to Fill Equations
topic Computation and Language
url https://arxiv.org/abs/2506.03763