Navigate the Unknown: Enhancing LLM Reasoning with Intrinsic Motivation Guided Exploration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Jingtong, Pan, Ling, Wang, Yejing, Zhong, Rui, Lu, Chi, Wang, Maolin, Cai, Qingpeng, Jiang, Peng, Zhao, Xiangyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908799328059392
author Gao, Jingtong
Pan, Ling
Wang, Yejing
Zhong, Rui
Lu, Chi
Wang, Maolin
Cai, Qingpeng
Jiang, Peng
Zhao, Xiangyu
author_facet Gao, Jingtong
Pan, Ling
Wang, Yejing
Zhong, Rui
Lu, Chi
Wang, Maolin
Cai, Qingpeng
Jiang, Peng
Zhao, Xiangyu
contents Reinforcement Learning (RL) has become a key approach for enhancing the reasoning capabilities of large language models. However, prevalent RL approaches like proximal policy optimization and group relative policy optimization suffer from sparse, outcome-based rewards and weak exploration incentives, limiting their effectiveness. Specifically, sparse rewards offer limited feedback, especially on difficult problems, and introduce biases favoring familiar trajectories over novel reasoning paths. These issues critically undermine performance on complex tasks that inherently require iterative reasoning. To overcome these challenges, we propose Intrinsic MotivAtion Guided exploratIoN for Enhanced reasoning (IMAGINE), which delivers dense rewards and encourages exploration. IMAGINE introduces three innovations: a trajectory-aware exploration reward that reduces token-level bias efficiently; an error-conditioned reward allocation that promotes efficient exploration on hard samples while stabilizing training; and an advantage-preserving integration mechanism that retains distributional integrity during learning. Experiments on four public datasets show that IMAGINE improves performance by 22.23% on AIME 2024.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17621
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Navigate the Unknown: Enhancing LLM Reasoning with Intrinsic Motivation Guided Exploration
Gao, Jingtong
Pan, Ling
Wang, Yejing
Zhong, Rui
Lu, Chi
Wang, Maolin
Cai, Qingpeng
Jiang, Peng
Zhao, Xiangyu
Machine Learning
Reinforcement Learning (RL) has become a key approach for enhancing the reasoning capabilities of large language models. However, prevalent RL approaches like proximal policy optimization and group relative policy optimization suffer from sparse, outcome-based rewards and weak exploration incentives, limiting their effectiveness. Specifically, sparse rewards offer limited feedback, especially on difficult problems, and introduce biases favoring familiar trajectories over novel reasoning paths. These issues critically undermine performance on complex tasks that inherently require iterative reasoning. To overcome these challenges, we propose Intrinsic MotivAtion Guided exploratIoN for Enhanced reasoning (IMAGINE), which delivers dense rewards and encourages exploration. IMAGINE introduces three innovations: a trajectory-aware exploration reward that reduces token-level bias efficiently; an error-conditioned reward allocation that promotes efficient exploration on hard samples while stabilizing training; and an advantage-preserving integration mechanism that retains distributional integrity during learning. Experiments on four public datasets show that IMAGINE improves performance by 22.23% on AIME 2024.
title Navigate the Unknown: Enhancing LLM Reasoning with Intrinsic Motivation Guided Exploration
topic Machine Learning
url https://arxiv.org/abs/2505.17621