Saved in:
Bibliographic Details
Main Authors: Wong, Zhen Hao, Deng, Jingwen, He, Runming, Chen, Zirong, You, Qijie, Dong, Hejun, Liang, Hao, Shen, Chengyu, Cui, Bin, Zhang, Wentao
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.04821
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913877736816640
author Wong, Zhen Hao
Deng, Jingwen
He, Runming
Chen, Zirong
You, Qijie
Dong, Hejun
Liang, Hao
Shen, Chengyu
Cui, Bin
Zhang, Wentao
author_facet Wong, Zhen Hao
Deng, Jingwen
He, Runming
Chen, Zirong
You, Qijie
Dong, Hejun
Liang, Hao
Shen, Chengyu
Cui, Bin
Zhang, Wentao
contents Large language models (LLMs) excel at many supervised tasks but often struggle with structured reasoning in unfamiliar settings. This discrepancy suggests that standard fine-tuning pipelines may instill narrow, domain-specific heuristics rather than fostering general-purpose thinking strategies. In this work, we propose a "play to learn" framework that fine-tunes LLMs through reinforcement learning on a suite of seven custom logic puzzles, each designed to cultivate distinct reasoning skills such as constraint propagation, spatial consistency, and symbolic deduction. Using a reinforcement learning setup with verifiable rewards, models receive binary feedback based on puzzle correctness, encouraging iterative, hypothesis-driven problem solving. We demonstrate that this training approach significantly improves out-of-distribution performance on a range of mathematical benchmarks, especially for mid-difficulty problems that require multi-step reasoning. Analyses across problem categories and difficulty levels reveal that puzzle training promotes transferable reasoning routines, strengthening algebraic manipulation, geometric inference, and combinatorial logic, while offering limited gains on rote or highly specialized tasks. These findings show that reinforcement learning over logic puzzles reshapes the internal reasoning of LLMs, enabling more robust and compositional generalization without relying on task-specific symbolic tools.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04821
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning
Wong, Zhen Hao
Deng, Jingwen
He, Runming
Chen, Zirong
You, Qijie
Dong, Hejun
Liang, Hao
Shen, Chengyu
Cui, Bin
Zhang, Wentao
Machine Learning
Large language models (LLMs) excel at many supervised tasks but often struggle with structured reasoning in unfamiliar settings. This discrepancy suggests that standard fine-tuning pipelines may instill narrow, domain-specific heuristics rather than fostering general-purpose thinking strategies. In this work, we propose a "play to learn" framework that fine-tunes LLMs through reinforcement learning on a suite of seven custom logic puzzles, each designed to cultivate distinct reasoning skills such as constraint propagation, spatial consistency, and symbolic deduction. Using a reinforcement learning setup with verifiable rewards, models receive binary feedback based on puzzle correctness, encouraging iterative, hypothesis-driven problem solving. We demonstrate that this training approach significantly improves out-of-distribution performance on a range of mathematical benchmarks, especially for mid-difficulty problems that require multi-step reasoning. Analyses across problem categories and difficulty levels reveal that puzzle training promotes transferable reasoning routines, strengthening algebraic manipulation, geometric inference, and combinatorial logic, while offering limited gains on rote or highly specialized tasks. These findings show that reinforcement learning over logic puzzles reshapes the internal reasoning of LLMs, enabling more robust and compositional generalization without relying on task-specific symbolic tools.
title LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2506.04821