Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Qiang, Song, Wuganjing, Lin, Zhenzhou, Chen, Feifan, Cai, Qiaolong, Li, Chen, Sui, Yongduo
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912671687770112
author Liu, Qiang
Song, Wuganjing
Lin, Zhenzhou
Chen, Feifan
Cai, Qiaolong
Li, Chen
Sui, Yongduo
author_facet Liu, Qiang
Song, Wuganjing
Lin, Zhenzhou
Chen, Feifan
Cai, Qiaolong
Li, Chen
Sui, Yongduo
contents The reasoning capabilities of Large Language Models (LLMs) are typically developed through the single-turn reinforcement learning, whereas real-world applications often involve multi-turn interactions with human feedback, leading to a potential mismatch between training and deployment conditions. In this work, we study whether multi-turn training with human feedback is necessary for reasoning tasks. We compare conventional single-turn training with three multi-turn strategies and reach contrary conclusions to previous research. We find that models trained in a single-turn setting generalize effectively to both single- and multi-turn evaluations, while models trained with multi-turn strategies exhibit a significant degradation in single-turn reasoning performance. These results suggest that for tasks with complete information, robust single-turn training remains more effective and reliable, as multi-turn training with basic feedback provides limited benefits and can even degrade reasoning capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21339
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning
Liu, Qiang
Song, Wuganjing
Lin, Zhenzhou
Chen, Feifan
Cai, Qiaolong
Li, Chen
Sui, Yongduo
Computation and Language
Information Theory
Machine Learning
The reasoning capabilities of Large Language Models (LLMs) are typically developed through the single-turn reinforcement learning, whereas real-world applications often involve multi-turn interactions with human feedback, leading to a potential mismatch between training and deployment conditions. In this work, we study whether multi-turn training with human feedback is necessary for reasoning tasks. We compare conventional single-turn training with three multi-turn strategies and reach contrary conclusions to previous research. We find that models trained in a single-turn setting generalize effectively to both single- and multi-turn evaluations, while models trained with multi-turn strategies exhibit a significant degradation in single-turn reasoning performance. These results suggest that for tasks with complete information, robust single-turn training remains more effective and reliable, as multi-turn training with basic feedback provides limited benefits and can even degrade reasoning capabilities.
title Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning
topic Computation and Language
Information Theory
Machine Learning
url https://arxiv.org/abs/2510.21339