Online Experiential Learning for Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Tianzhu, Dong, Li, Dong, Qingxiu, Wu, Xun, Huang, Shaohan, Wei, Furu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915870120345600
author Ye, Tianzhu
Dong, Li
Dong, Qingxiu
Wu, Xun
Huang, Shaohan
Wei, Furu
author_facet Ye, Tianzhu
Dong, Li
Dong, Qingxiu
Wu, Xun
Huang, Shaohan
Wei, Furu
contents The prevailing paradigm for improving large language models relies on offline training with human annotations or simulated environments, leaving the rich experience accumulated during real-world deployment entirely unexploited. We propose Online Experiential Learning (OEL), a framework that enables language models to continuously improve from their own deployment experience. OEL operates in two stages: first, transferable experiential knowledge is extracted and accumulated from interaction trajectories collected on the user side; second, this knowledge is consolidated into model parameters via on-policy context distillation, requiring no access to the user-side environment. The two stages are iterated to form an online learning loop, where the improved model collects higher-quality trajectories that yield richer experiential knowledge for subsequent rounds. We evaluate OEL on text-based game environments across multiple model scales and both thinking and non-thinking variants. OEL achieves consistent improvements over successive iterations, enhancing both task accuracy and token efficiency while preserving out-of-distribution performance. Our analysis further shows that extracted experiential knowledge is significantly more effective than raw trajectories, and that on-policy consistency between the knowledge source and the policy model is critical for effective learning.
format Preprint
id arxiv_https___arxiv_org_abs_2603_16856
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Online Experiential Learning for Language Models
Ye, Tianzhu
Dong, Li
Dong, Qingxiu
Wu, Xun
Huang, Shaohan
Wei, Furu
Computation and Language
The prevailing paradigm for improving large language models relies on offline training with human annotations or simulated environments, leaving the rich experience accumulated during real-world deployment entirely unexploited. We propose Online Experiential Learning (OEL), a framework that enables language models to continuously improve from their own deployment experience. OEL operates in two stages: first, transferable experiential knowledge is extracted and accumulated from interaction trajectories collected on the user side; second, this knowledge is consolidated into model parameters via on-policy context distillation, requiring no access to the user-side environment. The two stages are iterated to form an online learning loop, where the improved model collects higher-quality trajectories that yield richer experiential knowledge for subsequent rounds. We evaluate OEL on text-based game environments across multiple model scales and both thinking and non-thinking variants. OEL achieves consistent improvements over successive iterations, enhancing both task accuracy and token efficiency while preserving out-of-distribution performance. Our analysis further shows that extracted experiential knowledge is significantly more effective than raw trajectories, and that on-policy consistency between the knowledge source and the policy model is critical for effective learning.
title Online Experiential Learning for Language Models
topic Computation and Language
url https://arxiv.org/abs/2603.16856