Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Junsong, Zhou, Jie, Yang, Yutao, Zhan, Bihao, Pan, Qianjun, Ding, Yuyang, Chen, Qin, Bo, Jiang, Lin, Xin, He, Liang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908280731729920
author Li, Junsong
Zhou, Jie
Yang, Yutao
Zhan, Bihao
Pan, Qianjun
Ding, Yuyang
Chen, Qin
Bo, Jiang
Lin, Xin
He, Liang
author_facet Li, Junsong
Zhou, Jie
Yang, Yutao
Zhan, Bihao
Pan, Qianjun
Ding, Yuyang
Chen, Qin
Bo, Jiang
Lin, Xin
He, Liang
contents Automatic math correction aims to check students' solutions to mathematical problems via artificial intelligence technologies. Most existing studies focus on judging the final answer at the problem level, while they ignore detailed feedback on each step in a math problem-solving process, which requires abilities of semantic understanding and reasoning. In this paper, we propose a reinforcement learning (RL)-based method to boost large language model (LLM) for step-level automatic math correction, named StepAMC. Particularly, we convert the step-level automatic math correction within the text classification task into an RL problem to enhance the reasoning capabilities of LLMs. Then, we design a space-constrained policy network to improve the stability of RL. Then, we introduce a fine-grained reward network to convert the binary human feedback into a continuous value. We conduct extensive experiments over two benchmark datasets and the results show that our model outperforms the eleven strong baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18432
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning
Li, Junsong
Zhou, Jie
Yang, Yutao
Zhan, Bihao
Pan, Qianjun
Ding, Yuyang
Chen, Qin
Bo, Jiang
Lin, Xin
He, Liang
Computation and Language
Artificial Intelligence
Machine Learning
Automatic math correction aims to check students' solutions to mathematical problems via artificial intelligence technologies. Most existing studies focus on judging the final answer at the problem level, while they ignore detailed feedback on each step in a math problem-solving process, which requires abilities of semantic understanding and reasoning. In this paper, we propose a reinforcement learning (RL)-based method to boost large language model (LLM) for step-level automatic math correction, named StepAMC. Particularly, we convert the step-level automatic math correction within the text classification task into an RL problem to enhance the reasoning capabilities of LLMs. Then, we design a space-constrained policy network to improve the stability of RL. Then, we introduce a fine-grained reward network to convert the binary human feedback into a continuous value. We conduct extensive experiments over two benchmark datasets and the results show that our model outperforms the eleven strong baselines.
title Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.18432