Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Ang, Yuan, Zhihang, Zhang, Yang, Liu, Shouda, Wang, Yisen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909760854425600
author Li, Ang
Yuan, Zhihang
Zhang, Yang
Liu, Shouda
Wang, Yisen
author_facet Li, Ang
Yuan, Zhihang
Zhang, Yang
Liu, Shouda
Wang, Yisen
contents Reinforcement Learning with Verifiable Feedback (RLVF) has become a key technique for enhancing the reasoning abilities of Large Language Models (LLMs). However, its reliance on sparse, outcome based rewards, which only indicate if a final answer is correct or not, fails to provide granular guidance on the reasoning process itself. This limitation hinders efficient learning, as the model cannot distinguish between high quality and inefficient solutions, nor can it learn effectively from different types of failures. To address this, we observe that an LLMs self-certainty often correlates with task difficulty and solution quality. We introduce Difficulty Aware Certainty guided Exploration (DACE), a novel RL algorithm that leverages this insight to dynamically balance the exploration exploitation trade-off. DACE assesses task difficulty online based on the policys success rate. It then uses this signal to modulate an intrinsic reward: for difficult tasks where the model is struggling, DACE encourages exploration by penalizing high certainty; for easier tasks, it encourages learning efficiency by rewarding high certainty. Experiments on challenging mathematical reasoning benchmarks (AIME, MATH) show that DACE significantly outperforms strong baselines. The DACE-trained models not only achieve higher accuracy but also demonstrate more robust performance when scaling test-time compute, validating that our adaptive approach fosters effective exploration without sacrificing precision.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00125
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning
Li, Ang
Yuan, Zhihang
Zhang, Yang
Liu, Shouda
Wang, Yisen
Artificial Intelligence
Reinforcement Learning with Verifiable Feedback (RLVF) has become a key technique for enhancing the reasoning abilities of Large Language Models (LLMs). However, its reliance on sparse, outcome based rewards, which only indicate if a final answer is correct or not, fails to provide granular guidance on the reasoning process itself. This limitation hinders efficient learning, as the model cannot distinguish between high quality and inefficient solutions, nor can it learn effectively from different types of failures. To address this, we observe that an LLMs self-certainty often correlates with task difficulty and solution quality. We introduce Difficulty Aware Certainty guided Exploration (DACE), a novel RL algorithm that leverages this insight to dynamically balance the exploration exploitation trade-off. DACE assesses task difficulty online based on the policys success rate. It then uses this signal to modulate an intrinsic reward: for difficult tasks where the model is struggling, DACE encourages exploration by penalizing high certainty; for easier tasks, it encourages learning efficiency by rewarding high certainty. Experiments on challenging mathematical reasoning benchmarks (AIME, MATH) show that DACE significantly outperforms strong baselines. The DACE-trained models not only achieve higher accuracy but also demonstrate more robust performance when scaling test-time compute, validating that our adaptive approach fosters effective exploration without sacrificing precision.
title Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning
topic Artificial Intelligence
url https://arxiv.org/abs/2509.00125