Uncertainty-Guided Checkpoint Selection for Reinforcement Finetuning of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Manh, Nguyen, Dung, Do, Dai, Venkatesh, Svetha, Le, Hung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909900350685184
author Nguyen, Manh
Nguyen, Dung
Do, Dai
Venkatesh, Svetha
Le, Hung
author_facet Nguyen, Manh
Nguyen, Dung
Do, Dai
Venkatesh, Svetha
Le, Hung
contents Reinforcement learning (RL) finetuning is crucial to aligning large language models (LLMs), but the process is notoriously unstable and exhibits high variance across model checkpoints. In practice, selecting the best checkpoint is challenging: evaluating checkpoints on the validation set during training is computationally expensive and requires a good validation set, while relying on the final checkpoint provides no guarantee of good performance. We introduce an uncertainty-guided approach for checkpoint selection (UGCS) that avoids these pitfalls. Our method identifies hard question-answer pairs using per-sample uncertainty and ranks checkpoints by how well they handle these challenging cases. By averaging the rewards of the top-uncertain samples over a short training window, our method produces a stable and discriminative signal without additional forward passes or significant computation overhead. Experiments across three datasets and three LLMs demonstrate that it consistently identifies checkpoints with stronger generalization, outperforming traditional strategies such as relying on training or validation performance. These results highlight that models solving their hardest tasks with low uncertainty are the most reliable overall.
format Preprint
id arxiv_https___arxiv_org_abs_2511_09864
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Uncertainty-Guided Checkpoint Selection for Reinforcement Finetuning of Large Language Models
Nguyen, Manh
Nguyen, Dung
Do, Dai
Venkatesh, Svetha
Le, Hung
Machine Learning
Reinforcement learning (RL) finetuning is crucial to aligning large language models (LLMs), but the process is notoriously unstable and exhibits high variance across model checkpoints. In practice, selecting the best checkpoint is challenging: evaluating checkpoints on the validation set during training is computationally expensive and requires a good validation set, while relying on the final checkpoint provides no guarantee of good performance. We introduce an uncertainty-guided approach for checkpoint selection (UGCS) that avoids these pitfalls. Our method identifies hard question-answer pairs using per-sample uncertainty and ranks checkpoints by how well they handle these challenging cases. By averaging the rewards of the top-uncertain samples over a short training window, our method produces a stable and discriminative signal without additional forward passes or significant computation overhead. Experiments across three datasets and three LLMs demonstrate that it consistently identifies checkpoints with stronger generalization, outperforming traditional strategies such as relying on training or validation performance. These results highlight that models solving their hardest tasks with low uncertainty are the most reliable overall.
title Uncertainty-Guided Checkpoint Selection for Reinforcement Finetuning of Large Language Models
topic Machine Learning
url https://arxiv.org/abs/2511.09864