Risk-Sensitive RL for Alleviating Exploration Dilemmas in Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jiang, Yuhua, Huang, Jiawei, Yuan, Yufeng, Mao, Xin, Yue, Yu, Zhao, Qianchuan, Yan, Lin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915521102872576
author Jiang, Yuhua
Huang, Jiawei
Yuan, Yufeng
Mao, Xin
Yue, Yu
Zhao, Qianchuan
Yan, Lin
author_facet Jiang, Yuhua
Huang, Jiawei
Yuan, Yufeng
Mao, Xin
Yue, Yu
Zhao, Qianchuan
Yan, Lin
contents Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for enhancing Large Language Models (LLMs) on complex reasoning tasks. However, existing methods suffer from an exploration dilemma: the sharply peaked initial policies of pre-trained LLMs confine standard RL algorithms to a narrow set of solutions, boosting single-solution accuracy (pass@1) but suppressing solution diversity and multi-solution performance (pass@k). As a result, RLVR often distills existing capabilities rather than discovering new reasoning strategies. To overcome this, we introduce a Risk-Sensitive Reinforcement Learning framework. Our approach employs a risk-seeking objective that interpolates between mean and maximum rewards, leading to a novel algorithm, Risk-Sensitive GRPO (RS-GRPO), which drives deeper exploration by amplifying learning from challenging prompts. Remarkably, RS-GRPO is simple to implement, requiring only minor code modifications. On six mathematical reasoning benchmarks and with five different LLMs, RS-GRPO consistently improves pass@k performance while maintaining or enhancing pass@1 accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24261
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Risk-Sensitive RL for Alleviating Exploration Dilemmas in Large Language Models
Jiang, Yuhua
Huang, Jiawei
Yuan, Yufeng
Mao, Xin
Yue, Yu
Zhao, Qianchuan
Yan, Lin
Artificial Intelligence
Machine Learning
Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for enhancing Large Language Models (LLMs) on complex reasoning tasks. However, existing methods suffer from an exploration dilemma: the sharply peaked initial policies of pre-trained LLMs confine standard RL algorithms to a narrow set of solutions, boosting single-solution accuracy (pass@1) but suppressing solution diversity and multi-solution performance (pass@k). As a result, RLVR often distills existing capabilities rather than discovering new reasoning strategies. To overcome this, we introduce a Risk-Sensitive Reinforcement Learning framework. Our approach employs a risk-seeking objective that interpolates between mean and maximum rewards, leading to a novel algorithm, Risk-Sensitive GRPO (RS-GRPO), which drives deeper exploration by amplifying learning from challenging prompts. Remarkably, RS-GRPO is simple to implement, requiring only minor code modifications. On six mathematical reasoning benchmarks and with five different LLMs, RS-GRPO consistently improves pass@k performance while maintaining or enhancing pass@1 accuracy.
title Risk-Sensitive RL for Alleviating Exploration Dilemmas in Large Language Models
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.24261