Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Huang, Lei, Cheng, Xiang, Zhao, Chenxiao, Shen, Guobin, Yang, Junjie, Feng, Xiaocheng, Gu, Yuxuan, Yu, Xing, Qin, Bing
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908867097526272
author Huang, Lei
Cheng, Xiang
Zhao, Chenxiao
Shen, Guobin
Yang, Junjie
Feng, Xiaocheng
Gu, Yuxuan
Yu, Xing
Qin, Bing
author_facet Huang, Lei
Cheng, Xiang
Zhao, Chenxiao
Shen, Guobin
Yang, Junjie
Feng, Xiaocheng
Gu, Yuxuan
Yu, Xing
Qin, Bing
contents Large language models (LLMs) typically receive diverse natural language (NL) feedback through interaction with the environment. However, current reinforcement learning (RL) algorithms rely solely on scalar rewards, leaving the rich information in NL feedback underutilized and leading to inefficient exploration. In this work, we propose GOLF, an RL framework that explicitly exploits group-level language feedback to guide targeted exploration through actionable refinements. GOLF aggregates two complementary feedback sources: (i) external critiques that pinpoint errors or propose targeted fixes, and (ii) intra-group attempts that supply alternative partial ideas and diverse failure patterns. These group-level feedbacks are aggregated to produce high-quality refinements, which are adaptively injected into training as off-policy scaffolds to provide targeted guidance in sparse-reward regions. Meanwhile, GOLF jointly optimizes generation and refinement within a unified RL loop, creating a virtuous cycle that continuously improves both capabilities. Experiments on both verifiable and non-verifiable benchmarks show that GOLF achieves superior performance and exploration efficiency, achieving 2.2$\times$ improvements in sample efficiency compared to RL methods trained solely on scalar rewards. Code is available at https://github.com/LuckyyySTA/GOLF.
format Preprint
id arxiv_https___arxiv_org_abs_2603_04597
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning
Huang, Lei
Cheng, Xiang
Zhao, Chenxiao
Shen, Guobin
Yang, Junjie
Feng, Xiaocheng
Gu, Yuxuan
Yu, Xing
Qin, Bing
Computation and Language
Artificial Intelligence
Large language models (LLMs) typically receive diverse natural language (NL) feedback through interaction with the environment. However, current reinforcement learning (RL) algorithms rely solely on scalar rewards, leaving the rich information in NL feedback underutilized and leading to inefficient exploration. In this work, we propose GOLF, an RL framework that explicitly exploits group-level language feedback to guide targeted exploration through actionable refinements. GOLF aggregates two complementary feedback sources: (i) external critiques that pinpoint errors or propose targeted fixes, and (ii) intra-group attempts that supply alternative partial ideas and diverse failure patterns. These group-level feedbacks are aggregated to produce high-quality refinements, which are adaptively injected into training as off-policy scaffolds to provide targeted guidance in sparse-reward regions. Meanwhile, GOLF jointly optimizes generation and refinement within a unified RL loop, creating a virtuous cycle that continuously improves both capabilities. Experiments on both verifiable and non-verifiable benchmarks show that GOLF achieves superior performance and exploration efficiency, achieving 2.2$\times$ improvements in sample efficiency compared to RL methods trained solely on scalar rewards. Code is available at https://github.com/LuckyyySTA/GOLF.
title Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2603.04597