Scaling Self-Play with Self-Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Bailey, Luke, Wen, Kaiyue, Dong, Kefan, Hashimoto, Tatsunori, Ma, Tengyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STP: Self-play LLM Theorem Provers with Iterative Conjecturing and Proving
by: Dong, Kefan, et al.
Published: (2025)
by: Dong, Kefan, et al.
Published: (2025)
Configuration-to-Performance Scaling Law with Neural Ansatz
by: Zhang, Huaqing, et al.
Published: (2026)
by: Zhang, Huaqing, et al.
Published: (2026)
Pseudo-Formalization for Automatic Proof Verification
by: Barkallah, Slim, et al.
Published: (2026)
by: Barkallah, Slim, et al.
Published: (2026)
Formal Theorem Proving by Rewarding LLMs to Decompose Proofs Hierarchically
by: Dong, Kefan, et al.
Published: (2024)
by: Dong, Kefan, et al.
Published: (2024)
Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning
by: Mahankali, Arvind, et al.
Published: (2026)
by: Mahankali, Arvind, et al.
Published: (2026)
Linguistic Calibration of Long-Form Generations
by: Band, Neil, et al.
Published: (2024)
by: Band, Neil, et al.
Published: (2024)
Evaluating Self-Supervised Learning via Risk Decomposition
by: Dubois, Yann, et al.
Published: (2023)
by: Dubois, Yann, et al.
Published: (2023)
Fantastic Pretraining Optimizers and Where to Find Them
by: Wen, Kaiyue, et al.
Published: (2025)
by: Wen, Kaiyue, et al.
Published: (2025)
Scaling Laws for the Value of Individual Data Points in Machine Learning
by: Covert, Ian, et al.
Published: (2024)
by: Covert, Ian, et al.
Published: (2024)
Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective
by: Wen, Kaiyue, et al.
Published: (2024)
by: Wen, Kaiyue, et al.
Published: (2024)
Observational Scaling Laws and the Predictability of Language Model Performance
by: Ruan, Yangjun, et al.
Published: (2024)
by: Ruan, Yangjun, et al.
Published: (2024)
Language Models with Conformal Factuality Guarantees
by: Mohri, Christopher, et al.
Published: (2024)
by: Mohri, Christopher, et al.
Published: (2024)
A Bitter Lesson for Data Filtering
by: Mohri, Christopher, et al.
Published: (2026)
by: Mohri, Christopher, et al.
Published: (2026)
Understanding Finetuning for Factual Knowledge Extraction
by: Ghosal, Gaurav, et al.
Published: (2024)
by: Ghosal, Gaurav, et al.
Published: (2024)
Improving Pretraining Data Using Perplexity Correlations
by: Thrush, Tristan, et al.
Published: (2024)
by: Thrush, Tristan, et al.
Published: (2024)
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
by: Si, Chenglei, et al.
Published: (2024)
by: Si, Chenglei, et al.
Published: (2024)
Pre-training under infinite compute
by: Kim, Konwoo, et al.
Published: (2025)
by: Kim, Konwoo, et al.
Published: (2025)
Synthetic Data for any Differentiable Target
by: Thrush, Tristan, et al.
Published: (2026)
by: Thrush, Tristan, et al.
Published: (2026)
Self-Improving AI Agents through Self-Play
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning
by: Li, Shangzhe, et al.
Published: (2026)
by: Li, Shangzhe, et al.
Published: (2026)
Investigating Regularization of Self-Play Language Models
by: Alami, Reda, et al.
Published: (2024)
by: Alami, Reda, et al.
Published: (2024)
Stochastic Amortization: A Unified Approach to Accelerate Feature and Data Attribution
by: Covert, Ian, et al.
Published: (2024)
by: Covert, Ian, et al.
Published: (2024)
$π$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data
by: Zhang, Yaocheng, et al.
Published: (2026)
by: Zhang, Yaocheng, et al.
Published: (2026)
GASP: Guided Asymmetric Self-Play For Coding LLMs
by: Jana, Swadesh, et al.
Published: (2026)
by: Jana, Swadesh, et al.
Published: (2026)
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
by: Si, Chenglei, et al.
Published: (2025)
by: Si, Chenglei, et al.
Published: (2025)
Robust Distortion-free Watermarks for Language Models
by: Kuditipudi, Rohith, et al.
Published: (2023)
by: Kuditipudi, Rohith, et al.
Published: (2023)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
by: Tan, Zhewen, et al.
Published: (2026)
by: Tan, Zhewen, et al.
Published: (2026)
Diffusion Self-Weighted Guidance for Offline Reinforcement Learning
by: Tagle, Augusto, et al.
Published: (2025)
by: Tagle, Augusto, et al.
Published: (2025)
Non-Asymptotic Length Generalization
by: Chen, Thomas, et al.
Published: (2025)
by: Chen, Thomas, et al.
Published: (2025)
Momentum Guidance: Plug-and-Play Guidance for Flow Models
by: Liao, Runlong, et al.
Published: (2026)
by: Liao, Runlong, et al.
Published: (2026)
Offline Fictitious Self-Play for Competitive Games
by: Chen, Jingxiao, et al.
Published: (2024)
by: Chen, Jingxiao, et al.
Published: (2024)
A Theoretical Framework for Self-Play Theorem Proving Algorithms
by: Chen, Thomas, et al.
Published: (2026)
by: Chen, Thomas, et al.
Published: (2026)
Putting It All into Context: Simplifying Agents with LCLMs
by: Jiang, Mingjian, et al.
Published: (2025)
by: Jiang, Mingjian, et al.
Published: (2025)
Data-efficient pre-training by scaling synthetic megadocs
by: Kim, Konwoo, et al.
Published: (2026)
by: Kim, Konwoo, et al.
Published: (2026)
On the Learnability of Watermarks for Language Models
by: Gu, Chenchen, et al.
Published: (2023)
by: Gu, Chenchen, et al.
Published: (2023)
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
by: Dubois, Yann, et al.
Published: (2024)
by: Dubois, Yann, et al.
Published: (2024)
Reasoning to Learn from Latent Thoughts
by: Ruan, Yangjun, et al.
Published: (2025)
by: Ruan, Yangjun, et al.
Published: (2025)
Robust Autonomy Emerges from Self-Play
by: Cusumano-Towner, Marco, et al.
Published: (2025)
by: Cusumano-Towner, Marco, et al.
Published: (2025)
Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
by: Wang, Ru, et al.
Published: (2025)
by: Wang, Ru, et al.
Published: (2025)
R-Diverse: Mitigating Diversity Illusion in Self-Play LLM Training
by: Li, Gengsheng, et al.
Published: (2026)
by: Li, Gengsheng, et al.
Published: (2026)
Similar Items
-
STP: Self-play LLM Theorem Provers with Iterative Conjecturing and Proving
by: Dong, Kefan, et al.
Published: (2025) -
Configuration-to-Performance Scaling Law with Neural Ansatz
by: Zhang, Huaqing, et al.
Published: (2026) -
Pseudo-Formalization for Automatic Proof Verification
by: Barkallah, Slim, et al.
Published: (2026) -
Formal Theorem Proving by Rewarding LLMs to Decompose Proofs Hierarchically
by: Dong, Kefan, et al.
Published: (2024) -
Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning
by: Mahankali, Arvind, et al.
Published: (2026)