Language Self-Play For Data-Free Training
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914208744996864 |
|---|---|
| author | Kuba, Jakub Grudzien Gu, Mengting Ma, Qi Tian, Yuandong Mohan, Vijai Chen, Jason |
| author_facet | Kuba, Jakub Grudzien Gu, Mengting Ma, Qi Tian, Yuandong Mohan, Vijai Chen, Jason |
| contents | Large language models (LLMs) have advanced rapidly in recent years, driven by scale, abundant high-quality training data, and reinforcement learning. Yet this progress faces a fundamental bottleneck: the need for ever more data from which models can continue to learn. In this work, we propose a reinforcement learning approach that removes this dependency by enabling models to improve without additional data. Our method leverages a game-theoretic framework of self-play, where a model's capabilities are cast as performance in a competitive game and stronger policies emerge by having the model play against itself-a process we call Language Self-Play (LSP). Experiments with Llama-3.2-3B-Instruct on instruction-following, mathematics, and coding benchmarks show that pretrained models can be effectively improved with self-play alone. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_07414 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Language Self-Play For Data-Free Training Kuba, Jakub Grudzien Gu, Mengting Ma, Qi Tian, Yuandong Mohan, Vijai Chen, Jason Artificial Intelligence Computation and Language Computer Science and Game Theory Large language models (LLMs) have advanced rapidly in recent years, driven by scale, abundant high-quality training data, and reinforcement learning. Yet this progress faces a fundamental bottleneck: the need for ever more data from which models can continue to learn. In this work, we propose a reinforcement learning approach that removes this dependency by enabling models to improve without additional data. Our method leverages a game-theoretic framework of self-play, where a model's capabilities are cast as performance in a competitive game and stronger policies emerge by having the model play against itself-a process we call Language Self-Play (LSP). Experiments with Llama-3.2-3B-Instruct on instruction-following, mathematics, and coding benchmarks show that pretrained models can be effectively improved with self-play alone. |
| title | Language Self-Play For Data-Free Training |
| topic | Artificial Intelligence Computation and Language Computer Science and Game Theory |
| url | https://arxiv.org/abs/2509.07414 |