Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Xiao, Li, Zhongzhi, Gong, Yeyun, Shen, Yelong, Wu, Ying Nian, Guo, Zhijiang, Chen, Weizhu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
by: Liang, Xiao, et al.
Published: (2025)
by: Liang, Xiao, et al.
Published: (2025)
Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning
by: Huang, Yiming, et al.
Published: (2024)
by: Huang, Yiming, et al.
Published: (2024)
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
by: Gou, Zhibin, et al.
Published: (2023)
by: Gou, Zhibin, et al.
Published: (2023)
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
by: Gou, Zhibin, et al.
Published: (2023)
by: Gou, Zhibin, et al.
Published: (2023)
Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability
by: Liang, Xiao, et al.
Published: (2026)
by: Liang, Xiao, et al.
Published: (2026)
Exploring the Mystery of Influential Data for Mathematical Reasoning
by: Ni, Xinzhe, et al.
Published: (2024)
by: Ni, Xinzhe, et al.
Published: (2024)
Competition-Level Problems are Effective LLM Evaluators
by: Huang, Yiming, et al.
Published: (2023)
by: Huang, Yiming, et al.
Published: (2023)
Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation
by: Li, Zichong, et al.
Published: (2026)
by: Li, Zichong, et al.
Published: (2026)
Mixture of Neuron Experts
by: Cheng, Runxi, et al.
Published: (2025)
by: Cheng, Runxi, et al.
Published: (2025)
Rho-1: Not All Tokens Are What You Need
by: Lin, Zhenghao, et al.
Published: (2024)
by: Lin, Zhenghao, et al.
Published: (2024)
Enhancing Chain-of-Thoughts Prompting with Iterative Bootstrapping in Large Language Models
by: Sun, Jiashuo, et al.
Published: (2023)
by: Sun, Jiashuo, et al.
Published: (2023)
Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
by: Ren, Liliang, et al.
Published: (2024)
by: Ren, Liliang, et al.
Published: (2024)
Adapting LLM Agents with Universal Feedback in Communication
by: Wang, Kuan, et al.
Published: (2023)
by: Wang, Kuan, et al.
Published: (2023)
Gold-Medal-Level Olympiad Geometry Solving with Efficient Heuristic Auxiliary Constructions
by: Duan, Boyan, et al.
Published: (2025)
by: Duan, Boyan, et al.
Published: (2025)
Test-time Recursive Thinking: Self-Improvement without External Feedback
by: Zhuang, Yufan, et al.
Published: (2026)
by: Zhuang, Yufan, et al.
Published: (2026)
Efficient RLVR Training via Weighted Mutual Information Data Selection
by: Zhou, Xinyu, et al.
Published: (2026)
by: Zhou, Xinyu, et al.
Published: (2026)
Self-Distilled RLVR
by: Yang, Chenxu, et al.
Published: (2026)
by: Yang, Chenxu, et al.
Published: (2026)
Ensuring Safe and High-Quality Outputs: A Guideline Library Approach for Language Models
by: Luo, Yi, et al.
Published: (2024)
by: Luo, Yi, et al.
Published: (2024)
Training Matryoshka Mixture-of-Experts for Elastic Inference-Time Expert Utilization
by: Wang, Yaoxiang, et al.
Published: (2025)
by: Wang, Yaoxiang, et al.
Published: (2025)
Character-R1: Enhancing Role-Aware Reasoning in Role-Playing Agents via RLVR
by: Tang, Yihong, et al.
Published: (2026)
by: Tang, Yihong, et al.
Published: (2026)
Are LLMs Rigorous Logical Reasoners? Empowering Natural Language Proof Generation by Stepwise Decoding with Contrastive Learning
by: Su, Ying, et al.
Published: (2023)
by: Su, Ying, et al.
Published: (2023)
Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning
by: Park, Sungjin, et al.
Published: (2024)
by: Park, Sungjin, et al.
Published: (2024)
Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining Data
by: Deng, Haoran, et al.
Published: (2025)
by: Deng, Haoran, et al.
Published: (2025)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
by: Huang, Fanding, et al.
Published: (2025)
by: Huang, Fanding, et al.
Published: (2025)
Overcoming Vocabulary Mismatch: Vocabulary-agnostic Teacher Guided Language Modeling
by: Shin, Haebin, et al.
Published: (2025)
by: Shin, Haebin, et al.
Published: (2025)
Process-based Self-Rewarding Language Models
by: Zhang, Shimao, et al.
Published: (2025)
by: Zhang, Shimao, et al.
Published: (2025)
HiCaM: A Hierarchical-Causal Modification Framework for Long-Form Text Modification
by: Shi, Yuntao, et al.
Published: (2025)
by: Shi, Yuntao, et al.
Published: (2025)
AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators
by: He, Xingwei, et al.
Published: (2023)
by: He, Xingwei, et al.
Published: (2023)
Rethinking Language Model Scaling under Transferable Hypersphere Optimization
by: Ren, Liliang, et al.
Published: (2026)
by: Ren, Liliang, et al.
Published: (2026)
Improving Data and Reward Design for Scientific Reasoning in Large Language Models
by: Chen, Zijie, et al.
Published: (2026)
by: Chen, Zijie, et al.
Published: (2026)
Length-Unbiased Sequence Policy Optimization: Revealing and Controlling Response Length Variation in RLVR
by: Liu, Fanfan, et al.
Published: (2026)
by: Liu, Fanfan, et al.
Published: (2026)
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR
by: Lee, Chanuk, et al.
Published: (2026)
by: Lee, Chanuk, et al.
Published: (2026)
ThetaEvolve: Test-time Learning on Open Problems
by: Wang, Yiping, et al.
Published: (2025)
by: Wang, Yiping, et al.
Published: (2025)
The Unlearnability Phenomenon in RLVR for Language Models
by: Chen, Yulin, et al.
Published: (2026)
by: Chen, Yulin, et al.
Published: (2026)
Document Reconstruction Unlocks Scalable Long-Context RLVR
by: Xiao, Yao, et al.
Published: (2026)
by: Xiao, Yao, et al.
Published: (2026)
APOLLO: An Optimized Training Approach for Long-form Numerical Reasoning
by: Sun, Jiashuo, et al.
Published: (2022)
by: Sun, Jiashuo, et al.
Published: (2022)
Rewards as Labels: Revisiting RLVR from a Classification Perspective
by: Zhai, Zepeng, et al.
Published: (2026)
by: Zhai, Zepeng, et al.
Published: (2026)
Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains
by: Yuan, Zhonghang, et al.
Published: (2026)
by: Yuan, Zhonghang, et al.
Published: (2026)
Multi-LoRA Composition for Image Generation
by: Zhong, Ming, et al.
Published: (2024)
by: Zhong, Ming, et al.
Published: (2024)
When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond?
by: Zhou, Xinyu, et al.
Published: (2026)
by: Zhou, Xinyu, et al.
Published: (2026)
Similar Items
-
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
by: Liang, Xiao, et al.
Published: (2025) -
Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning
by: Huang, Yiming, et al.
Published: (2024) -
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
by: Gou, Zhibin, et al.
Published: (2023) -
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
by: Gou, Zhibin, et al.
Published: (2023) -
Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability
by: Liang, Xiao, et al.
Published: (2026)