MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Xuancheng, Li, Haitao, Zhou, Yujia, YiqunLiu, Ai, Qingyao
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916069815353344
author Li, Xuancheng
Li, Haitao
Zhou, Yujia
YiqunLiu
Ai, Qingyao
author_facet Li, Xuancheng
Li, Haitao
Zhou, Yujia
YiqunLiu
Ai, Qingyao
contents Reinforcement Learning with Verifiable Rewards (RLVR) is widely used to improve reasoning across domains, but outcome-only scalar rewards are often sparse and uninformative. This limitation is especially severe for failed samples, where scalar rewards indicate only that a solution is incorrect without explaining why the reasoning breaks down. In this paper, we leverage richer verbal feedback to guide RLVR on failed samples and convert feedback-induced progress into trainable learning signals. We propose MulFeRL (Multi-turn Feedback-guided Reinforcement Learning), a multi-turn, event-triggered RLVR framework that combines progress induction for feedback-guided regeneration of failed samples, progress credit assignment for learning from verifier-confirmed progress, and structured feedback injection for integrating feedback into the model's reasoning process. Trained on sampled OpenR1-Math, MulFeRL outperforms supervised, self-distillation-based, and RLVR baselines in-domain, while also showing strong out-of-domain generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22900
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop
Li, Xuancheng
Li, Haitao
Zhou, Yujia
YiqunLiu
Ai, Qingyao
Artificial Intelligence
Reinforcement Learning with Verifiable Rewards (RLVR) is widely used to improve reasoning across domains, but outcome-only scalar rewards are often sparse and uninformative. This limitation is especially severe for failed samples, where scalar rewards indicate only that a solution is incorrect without explaining why the reasoning breaks down. In this paper, we leverage richer verbal feedback to guide RLVR on failed samples and convert feedback-induced progress into trainable learning signals. We propose MulFeRL (Multi-turn Feedback-guided Reinforcement Learning), a multi-turn, event-triggered RLVR framework that combines progress induction for feedback-guided regeneration of failed samples, progress credit assignment for learning from verifier-confirmed progress, and structured feedback injection for integrating feedback into the model's reasoning process. Trained on sampled OpenR1-Math, MulFeRL outperforms supervised, self-distillation-based, and RLVR baselines in-domain, while also showing strong out-of-domain generalization.
title MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop
topic Artificial Intelligence
url https://arxiv.org/abs/2601.22900