The Era of Real-World Human Interaction: RL from User Conversations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Chuanyang, Xu, Jing, Liu, Bo, Tao, Leitian, Golovneva, Olga, Shu, Tianmin, Zhao, Wenting, Li, Xian, Weston, Jason
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918150774194176
author Jin, Chuanyang
Xu, Jing
Liu, Bo
Tao, Leitian
Golovneva, Olga
Shu, Tianmin
Zhao, Wenting
Li, Xian
Weston, Jason
author_facet Jin, Chuanyang
Xu, Jing
Liu, Bo
Tao, Leitian
Golovneva, Olga
Shu, Tianmin
Zhao, Wenting
Li, Xian
Weston, Jason
contents We posit that to achieve continual model improvement and multifaceted alignment, future models must learn from natural human interaction. Current conversational models are aligned using pre-annotated, expert-generated human feedback. In this work, we introduce Reinforcement Learning from Human Interaction (RLHI), a paradigm that learns directly from in-the-wild user conversations. We develop two complementary methods: (1) RLHI with User-Guided Rewrites, which revises unsatisfactory model outputs based on users' natural-language follow-up responses, (2) RLHI with User-Based Rewards, which learns via a reward model conditioned on knowledge of the user's long-term interaction history (termed persona). Together, these methods link long-term user personas to turn-level preferences via persona-conditioned preference optimization. Trained on conversations derived from WildChat, both RLHI variants outperform strong baselines in personalization and instruction-following, and similar feedback enhances performance on reasoning benchmarks. These results suggest organic human interaction offers scalable, effective supervision for personalized alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25137
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Era of Real-World Human Interaction: RL from User Conversations
Jin, Chuanyang
Xu, Jing
Liu, Bo
Tao, Leitian
Golovneva, Olga
Shu, Tianmin
Zhao, Wenting
Li, Xian
Weston, Jason
Artificial Intelligence
Computation and Language
Machine Learning
We posit that to achieve continual model improvement and multifaceted alignment, future models must learn from natural human interaction. Current conversational models are aligned using pre-annotated, expert-generated human feedback. In this work, we introduce Reinforcement Learning from Human Interaction (RLHI), a paradigm that learns directly from in-the-wild user conversations. We develop two complementary methods: (1) RLHI with User-Guided Rewrites, which revises unsatisfactory model outputs based on users' natural-language follow-up responses, (2) RLHI with User-Based Rewards, which learns via a reward model conditioned on knowledge of the user's long-term interaction history (termed persona). Together, these methods link long-term user personas to turn-level preferences via persona-conditioned preference optimization. Trained on conversations derived from WildChat, both RLHI variants outperform strong baselines in personalization and instruction-following, and similar feedback enhances performance on reasoning benchmarks. These results suggest organic human interaction offers scalable, effective supervision for personalized alignment.
title The Era of Real-World Human Interaction: RL from User Conversations
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.25137