Language Models Can Learn from Verbal Feedback Without Scalar Rewards

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Renjie, Liu, Zichen, Liu, Xiangyan, Du, Chao, Lin, Min, Chen, Wenhu, Lu, Wei, Pang, Tianyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909809729601536
author Luo, Renjie
Liu, Zichen
Liu, Xiangyan
Du, Chao
Lin, Min
Chen, Wenhu
Lu, Wei
Pang, Tianyu
author_facet Luo, Renjie
Liu, Zichen
Liu, Xiangyan
Du, Chao
Lin, Min
Chen, Wenhu
Lu, Wei
Pang, Tianyu
contents LLMs are often trained with RL from human or AI feedback, yet such methods typically compress nuanced feedback into scalar rewards, discarding much of their richness and inducing scale imbalance. We propose treating verbal feedback as a conditioning signal. Inspired by language priors in text-to-image generation, which enable novel outputs from unseen prompts, we introduce the feedback-conditional policy (FCP). FCP learns directly from response-feedback pairs, approximating the feedback-conditional posterior through maximum likelihood training on offline data. We further develop an online bootstrapping stage where the policy generates under positive conditions and receives fresh feedback to refine itself. This reframes feedback-driven learning as conditional generation rather than reward optimization, offering a more expressive way for LLMs to directly learn from verbal feedback. Our code is available at https://github.com/sail-sg/feedback-conditional-policy.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22638
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Language Models Can Learn from Verbal Feedback Without Scalar Rewards
Luo, Renjie
Liu, Zichen
Liu, Xiangyan
Du, Chao
Lin, Min
Chen, Wenhu
Lu, Wei
Pang, Tianyu
Computation and Language
Artificial Intelligence
Machine Learning
LLMs are often trained with RL from human or AI feedback, yet such methods typically compress nuanced feedback into scalar rewards, discarding much of their richness and inducing scale imbalance. We propose treating verbal feedback as a conditioning signal. Inspired by language priors in text-to-image generation, which enable novel outputs from unseen prompts, we introduce the feedback-conditional policy (FCP). FCP learns directly from response-feedback pairs, approximating the feedback-conditional posterior through maximum likelihood training on offline data. We further develop an online bootstrapping stage where the policy generates under positive conditions and receives fresh feedback to refine itself. This reframes feedback-driven learning as conditional generation rather than reward optimization, offering a more expressive way for LLMs to directly learn from verbal feedback. Our code is available at https://github.com/sail-sg/feedback-conditional-policy.
title Language Models Can Learn from Verbal Feedback Without Scalar Rewards
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.22638