Language Self-Play For Data-Free Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kuba, Jakub Grudzien, Gu, Mengting, Ma, Qi, Tian, Yuandong, Mohan, Vijai, Chen, Jason
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914208744996864
author Kuba, Jakub Grudzien
Gu, Mengting
Ma, Qi
Tian, Yuandong
Mohan, Vijai
Chen, Jason
author_facet Kuba, Jakub Grudzien
Gu, Mengting
Ma, Qi
Tian, Yuandong
Mohan, Vijai
Chen, Jason
contents Large language models (LLMs) have advanced rapidly in recent years, driven by scale, abundant high-quality training data, and reinforcement learning. Yet this progress faces a fundamental bottleneck: the need for ever more data from which models can continue to learn. In this work, we propose a reinforcement learning approach that removes this dependency by enabling models to improve without additional data. Our method leverages a game-theoretic framework of self-play, where a model's capabilities are cast as performance in a competitive game and stronger policies emerge by having the model play against itself-a process we call Language Self-Play (LSP). Experiments with Llama-3.2-3B-Instruct on instruction-following, mathematics, and coding benchmarks show that pretrained models can be effectively improved with self-play alone.
format Preprint
id arxiv_https___arxiv_org_abs_2509_07414
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Language Self-Play For Data-Free Training
Kuba, Jakub Grudzien
Gu, Mengting
Ma, Qi
Tian, Yuandong
Mohan, Vijai
Chen, Jason
Artificial Intelligence
Computation and Language
Computer Science and Game Theory
Large language models (LLMs) have advanced rapidly in recent years, driven by scale, abundant high-quality training data, and reinforcement learning. Yet this progress faces a fundamental bottleneck: the need for ever more data from which models can continue to learn. In this work, we propose a reinforcement learning approach that removes this dependency by enabling models to improve without additional data. Our method leverages a game-theoretic framework of self-play, where a model's capabilities are cast as performance in a competitive game and stronger policies emerge by having the model play against itself-a process we call Language Self-Play (LSP). Experiments with Llama-3.2-3B-Instruct on instruction-following, mathematics, and coding benchmarks show that pretrained models can be effectively improved with self-play alone.
title Language Self-Play For Data-Free Training
topic Artificial Intelligence
Computation and Language
Computer Science and Game Theory
url https://arxiv.org/abs/2509.07414