Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Junsong, Zhou, Jie, Yang, Yutao, Zhan, Bihao, Pan, Qianjun, Ding, Yuyang, Chen, Qin, Bo, Jiang, Lin, Xin, He, Liang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reinforced Interactive Continual Learning via Real-time Noisy Human Feedback
by: Yang, Yutao, et al.
Published: (2025)
by: Yang, Yutao, et al.
Published: (2025)
LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
by: Li, Junsong, et al.
Published: (2025)
by: Li, Junsong, et al.
Published: (2025)
Forget What's Sensitive, Remember What Matters: Token-Level Differential Privacy in Memory Sculpting for Continual Learning
by: Zhan, Bihao, et al.
Published: (2025)
by: Zhan, Bihao, et al.
Published: (2025)
AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
by: Yang, Yutao, et al.
Published: (2026)
by: Yang, Yutao, et al.
Published: (2026)
Black-box Model Merging for Language-Model-as-a-Service with Massive Model Repositories
by: Chen, Shilian, et al.
Published: (2025)
by: Chen, Shilian, et al.
Published: (2025)
A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law
by: Pan, Qianjun, et al.
Published: (2025)
by: Pan, Qianjun, et al.
Published: (2025)
CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models
by: Liu, Wentao, et al.
Published: (2024)
by: Liu, Wentao, et al.
Published: (2024)
PsychAgent: An Experience-Driven Lifelong Learning Agent for Self-Evolving Psychological Counselor
by: Yang, Yutao, et al.
Published: (2026)
by: Yang, Yutao, et al.
Published: (2026)
Boosting Large Language Models with Socratic Method for Conversational Mathematics Teaching
by: Ding, Yuyang, et al.
Published: (2024)
by: Ding, Yuyang, et al.
Published: (2024)
MathMistake Checker: A Comprehensive Demonstration for Step-by-Step Math Problem Mistake Finding by Prompt-Guided LLMs
by: Zhang, Tianyang, et al.
Published: (2025)
by: Zhang, Tianyang, et al.
Published: (2025)
PsychEval: A Multi-Session and Multi-Therapy Benchmark for High-Realism AI Psychological Counselor
by: Pan, Qianjun, et al.
Published: (2026)
by: Pan, Qianjun, et al.
Published: (2026)
Mathematical Language Models: A Survey
by: Liu, Wentao, et al.
Published: (2023)
by: Liu, Wentao, et al.
Published: (2023)
WirelessMathLM: Teaching Mathematical Reasoning for LLMs in Wireless Communications with Reinforcement Learning
by: Li, Xin, et al.
Published: (2025)
by: Li, Xin, et al.
Published: (2025)
SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese
by: Xu, Liang, et al.
Published: (2024)
by: Xu, Liang, et al.
Published: (2024)
Building Self-Evolving Agents via Experience-Driven Lifelong Learning: A Framework and Benchmark
by: Cai, Yuxuan, et al.
Published: (2025)
by: Cai, Yuxuan, et al.
Published: (2025)
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
by: Wang, Peiyi, et al.
Published: (2023)
by: Wang, Peiyi, et al.
Published: (2023)
Let's Verify Math Questions Step by Step
by: Shen, Chengyu, et al.
Published: (2025)
by: Shen, Chengyu, et al.
Published: (2025)
The CompMath-MCQ Dataset: Are LLMs Ready for Higher-Level Math?
by: Raimondi, Bianca, et al.
Published: (2026)
by: Raimondi, Bianca, et al.
Published: (2026)
A Step-by-step Introduction to the Implementation of Automatic Differentiation
by: Fang, Yu-Hsueh, et al.
Published: (2024)
by: Fang, Yu-Hsueh, et al.
Published: (2024)
WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning
by: Lu, Zimu, et al.
Published: (2025)
by: Lu, Zimu, et al.
Published: (2025)
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs
by: Lu, Zimu, et al.
Published: (2024)
by: Lu, Zimu, et al.
Published: (2024)
Efficient Excited-State Calculations for Molecules Based on Contextual Subspace Method and Symmetry Optimizations
by: Yao, Qianjun, et al.
Published: (2025)
by: Yao, Qianjun, et al.
Published: (2025)
LEMMA: Learning from Errors for MatheMatical Advancement in LLMs
by: Pan, Zhuoshi, et al.
Published: (2025)
by: Pan, Zhuoshi, et al.
Published: (2025)
Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems
by: Zhong, Qihuang, et al.
Published: (2024)
by: Zhong, Qihuang, et al.
Published: (2024)
Let's Rectify Step by Step: Improving Aspect-based Sentiment Analysis with Diffusion Models
by: Liu, Shunyu, et al.
Published: (2024)
by: Liu, Shunyu, et al.
Published: (2024)
Optimal Refund Mechanism with Consumer Learning
by: Lyu, Qianjun
Published: (2024)
by: Lyu, Qianjun
Published: (2024)
Optimal Refund Mechanism With Consumer Learning
by: Qianjun Lyu
Published: (2026)
by: Qianjun Lyu
Published: (2026)
MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code
by: Lu, Zimu, et al.
Published: (2024)
by: Lu, Zimu, et al.
Published: (2024)
Truncated Rectified Flow Policy for Reinforcement Learning with One-Step Sampling
by: Zhou, Xubin, et al.
Published: (2026)
by: Zhou, Xubin, et al.
Published: (2026)
Recent Advances of Foundation Language Models-based Continual Learning: A Survey
by: Yang, Yutao, et al.
Published: (2024)
by: Yang, Yutao, et al.
Published: (2024)
GRC-Net: Gram Residual Co-attention Net for epilepsy prediction
by: You, Bihao, et al.
Published: (2025)
by: You, Bihao, et al.
Published: (2025)
UCO: A Multi-Turn Interactive Reinforcement Learning Method for Adaptive Teaching with Large Language Models
by: Wei, Shouang, et al.
Published: (2025)
by: Wei, Shouang, et al.
Published: (2025)
StepMathAgent: A Step-Wise Agent for Evaluating Mathematical Processes through Tree-of-Error
by: Yang, Shu-Xun, et al.
Published: (2025)
by: Yang, Shu-Xun, et al.
Published: (2025)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
by: Pandit, Shrey, et al.
Published: (2025)
by: Pandit, Shrey, et al.
Published: (2025)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
by: Balunović, Mislav, et al.
Published: (2025)
by: Balunović, Mislav, et al.
Published: (2025)
MathFimer: Enhancing Mathematical Reasoning by Expanding Reasoning Steps through Fill-in-the-Middle Task
by: Yan, Yuchen, et al.
Published: (2025)
by: Yan, Yuchen, et al.
Published: (2025)
MARGE: Improving Math Reasoning for LLMs with Guided Exploration
by: Gao, Jingyue, et al.
Published: (2025)
by: Gao, Jingyue, et al.
Published: (2025)
SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency Distillation
by: Chen, Junsong, et al.
Published: (2025)
by: Chen, Junsong, et al.
Published: (2025)
ShadowCode: Towards (Automatic) External Prompt Injection Attack against Code LLMs
by: Yang, Yuchen, et al.
Published: (2024)
by: Yang, Yuchen, et al.
Published: (2024)
R3-RAG: Learning Step-by-Step Reasoning and Retrieval for LLMs via Reinforcement Learning
by: Li, Yuan, et al.
Published: (2025)
by: Li, Yuan, et al.
Published: (2025)
Similar Items
-
Reinforced Interactive Continual Learning via Real-time Noisy Human Feedback
by: Yang, Yutao, et al.
Published: (2025) -
LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
by: Li, Junsong, et al.
Published: (2025) -
Forget What's Sensitive, Remember What Matters: Token-Level Differential Privacy in Memory Sculpting for Continual Learning
by: Zhan, Bihao, et al.
Published: (2025) -
AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
by: Yang, Yutao, et al.
Published: (2026) -
Black-box Model Merging for Language-Model-as-a-Service with Massive Model Repositories
by: Chen, Shilian, et al.
Published: (2025)