A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Licheng, Wang, Zihan, Li, Linjie, Xu, Chenwei, Lu, Yiping, Liu, Han, Sil, Avirup, Li, Manling
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908498254626816
author Liu, Licheng
Wang, Zihan
Li, Linjie
Xu, Chenwei
Lu, Yiping
Liu, Han
Sil, Avirup
Li, Manling
author_facet Liu, Licheng
Wang, Zihan
Li, Linjie
Xu, Chenwei
Lu, Yiping
Liu, Han
Sil, Avirup
Li, Manling
contents Multi-turn problem solving is critical yet challenging for Large Reasoning Models (LRMs) to reflect on their reasoning and revise from feedback. Existing Reinforcement Learning (RL) methods train large reasoning models on a single-turn paradigm with verifiable rewards. However, we observe that models trained with existing RL paradigms often lose their ability to solve problems across multiple turns and struggle to revise answers based on contextual feedback, leading to repetitive responses. We ask: can LRMs learn to reflect their answers in a multi-turn context? In this work, we find that training models with multi-turn RL using only unary feedback (e.g., "Let's try again") after wrong answers can improve both single-turn performance and multi-turn reasoning. We introduce Unary Feedback as Observation (UFO) for reinforcement learning, which uses minimal yet common unary user feedback during iterative problem solving. It can be easily applied to existing single-turn RL training setups. Experimental results show that RL training with UFO keeps single-turn performance and improves multi-turn reasoning accuracy by up to 14%, enabling language models to better react to feedback in multi-turn problem solving. To further minimize the number of turns needed for a correct answer while encouraging diverse reasoning when mistakes occur, we design reward structures that guide models to produce careful and deliberate answers in each turn. Code: https://github.com/lichengliu03/unary-feedback
format Preprint
id arxiv_https___arxiv_org_abs_2507_14295
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning
Liu, Licheng
Wang, Zihan
Li, Linjie
Xu, Chenwei
Lu, Yiping
Liu, Han
Sil, Avirup
Li, Manling
Machine Learning
Artificial Intelligence
Multi-turn problem solving is critical yet challenging for Large Reasoning Models (LRMs) to reflect on their reasoning and revise from feedback. Existing Reinforcement Learning (RL) methods train large reasoning models on a single-turn paradigm with verifiable rewards. However, we observe that models trained with existing RL paradigms often lose their ability to solve problems across multiple turns and struggle to revise answers based on contextual feedback, leading to repetitive responses. We ask: can LRMs learn to reflect their answers in a multi-turn context? In this work, we find that training models with multi-turn RL using only unary feedback (e.g., "Let's try again") after wrong answers can improve both single-turn performance and multi-turn reasoning. We introduce Unary Feedback as Observation (UFO) for reinforcement learning, which uses minimal yet common unary user feedback during iterative problem solving. It can be easily applied to existing single-turn RL training setups. Experimental results show that RL training with UFO keeps single-turn performance and improves multi-turn reasoning accuracy by up to 14%, enabling language models to better react to feedback in multi-turn problem solving. To further minimize the number of turns needed for a correct answer while encouraging diverse reasoning when mistakes occur, we design reward structures that guide models to produce careful and deliberate answers in each turn. Code: https://github.com/lichengliu03/unary-feedback
title A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2507.14295