TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Jinyang, Liao, Chonghua, Feng, Mingkuan, Zhang, Shuai, Wen, Zhengqi, Luo, Haoran, Yang, Ling, Xu, Huazhe, Tao, Jianhua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913131232493568
author Wu, Jinyang
Liao, Chonghua
Feng, Mingkuan
Zhang, Shuai
Wen, Zhengqi
Luo, Haoran
Yang, Ling
Xu, Huazhe
Tao, Jianhua
author_facet Wu, Jinyang
Liao, Chonghua
Feng, Mingkuan
Zhang, Shuai
Wen, Zhengqi
Luo, Haoran
Yang, Ling
Xu, Huazhe
Tao, Jianhua
contents Reinforcement learning (RL) has emerged as an effective paradigm for enhancing model reasoning. However, existing RL methods like GRPO typically rely on unstructured self-sampling to fit scalar rewards, often producing inefficient rollouts that fail to capture transferable problem-solving strategies. To address this limitation, we propose **TemplateRL**, a structured template-guided RL framework that augments policy optimization with explicit template guidance. Our approach first constructs a problem-solving template library via MCTS on a small seed set, then seamlessly integrates this high-level structured guidance into RL training. By guiding rollout generation to align with proven template structures, TemplateRL significantly improves high-quality trajectory hit rates while reducing ineffective exploration. This structure-guided design steers the policy toward validated strategic patterns, stabilizing training dynamics, and enhancing RL sampling efficiency. Notably, the explicit template library is interpretable, editable, and supports online updates-enabling continuous updates during both training and inference. Extensive experiments demonstrate that TemplateRL outperforms GRPO by 99% on AIME and 41% on AMC, with superior stability on weak models and remarkable cross-domain generalization, highlighting its potential for broader tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15692
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning
Wu, Jinyang
Liao, Chonghua
Feng, Mingkuan
Zhang, Shuai
Wen, Zhengqi
Luo, Haoran
Yang, Ling
Xu, Huazhe
Tao, Jianhua
Computation and Language
Machine Learning
Reinforcement learning (RL) has emerged as an effective paradigm for enhancing model reasoning. However, existing RL methods like GRPO typically rely on unstructured self-sampling to fit scalar rewards, often producing inefficient rollouts that fail to capture transferable problem-solving strategies. To address this limitation, we propose **TemplateRL**, a structured template-guided RL framework that augments policy optimization with explicit template guidance. Our approach first constructs a problem-solving template library via MCTS on a small seed set, then seamlessly integrates this high-level structured guidance into RL training. By guiding rollout generation to align with proven template structures, TemplateRL significantly improves high-quality trajectory hit rates while reducing ineffective exploration. This structure-guided design steers the policy toward validated strategic patterns, stabilizing training dynamics, and enhancing RL sampling efficiency. Notably, the explicit template library is interpretable, editable, and supports online updates-enabling continuous updates during both training and inference. Extensive experiments demonstrate that TemplateRL outperforms GRPO by 99% on AIME and 41% on AMC, with superior stability on weak models and remarkable cross-domain generalization, highlighting its potential for broader tasks.
title TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.15692