SPELL: Self-Play Reinforcement Learning for Evolving Long-Context Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Ziyi, Shen, Weizhou, Li, Chenliang, Chen, Ruijun, Wan, Fanqi, Yan, Ming, Quan, Xiaojun, Huang, Fei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915858190696448
author Yang, Ziyi
Shen, Weizhou
Li, Chenliang
Chen, Ruijun
Wan, Fanqi
Yan, Ming
Quan, Xiaojun
Huang, Fei
author_facet Yang, Ziyi
Shen, Weizhou
Li, Chenliang
Chen, Ruijun
Wan, Fanqi
Yan, Ming
Quan, Xiaojun
Huang, Fei
contents Progress in long-context reasoning for large language models (LLMs) has lagged behind other recent advances. This gap arises not only from the intrinsic difficulty of processing long texts, but also from the scarcity of reliable human annotations and programmatically verifiable reward signals. In this paper, we propose SPELL, a multi-role self-play reinforcement learning framework that enables scalable, label-free optimization for long-context reasoning. SPELL integrates three cyclical roles-questioner, responder, and verifier-within a single model to enable continual self-improvement. The questioner generates questions from raw documents paired with reference answers; the responder learns to solve these questions based on the documents; and the verifier evaluates semantic equivalence between the responder's output and the questioner's reference answer, producing reward signals to guide continual training. To stabilize training, we introduce an automated curriculum that gradually increases document length and a reward function that adapts question difficulty to the model's evolving capabilities. Extensive experiments on six long-context benchmarks show that SPELL consistently improves performance across diverse LLMs and outperforms equally sized models fine-tuned on large-scale annotated data. Notably, SPELL achieves an average 7.6-point gain in pass@8 on the strong reasoning model Qwen3-30B-A3B-Thinking, raising its performance ceiling and showing promise for scaling to even more capable models. Our code is available at https://github.com/Tongyi-Zhiwen/Qwen-Doc.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23863
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SPELL: Self-Play Reinforcement Learning for Evolving Long-Context Language Models
Yang, Ziyi
Shen, Weizhou
Li, Chenliang
Chen, Ruijun
Wan, Fanqi
Yan, Ming
Quan, Xiaojun
Huang, Fei
Computation and Language
Progress in long-context reasoning for large language models (LLMs) has lagged behind other recent advances. This gap arises not only from the intrinsic difficulty of processing long texts, but also from the scarcity of reliable human annotations and programmatically verifiable reward signals. In this paper, we propose SPELL, a multi-role self-play reinforcement learning framework that enables scalable, label-free optimization for long-context reasoning. SPELL integrates three cyclical roles-questioner, responder, and verifier-within a single model to enable continual self-improvement. The questioner generates questions from raw documents paired with reference answers; the responder learns to solve these questions based on the documents; and the verifier evaluates semantic equivalence between the responder's output and the questioner's reference answer, producing reward signals to guide continual training. To stabilize training, we introduce an automated curriculum that gradually increases document length and a reward function that adapts question difficulty to the model's evolving capabilities. Extensive experiments on six long-context benchmarks show that SPELL consistently improves performance across diverse LLMs and outperforms equally sized models fine-tuned on large-scale annotated data. Notably, SPELL achieves an average 7.6-point gain in pass@8 on the strong reasoning model Qwen3-30B-A3B-Thinking, raising its performance ceiling and showing promise for scaling to even more capable models. Our code is available at https://github.com/Tongyi-Zhiwen/Qwen-Doc.
title SPELL: Self-Play Reinforcement Learning for Evolving Long-Context Language Models
topic Computation and Language
url https://arxiv.org/abs/2509.23863