Predictive Scaling Laws for Efficient GRPO Training of Large Reasoning Models
Fuente:
arXiv
Saved in:
| Main Authors: | Nimmaturi, Datta, Bhargava, Vaishnavi, Ghosh, Rajat, George, Johnu, Dutta, Debojyoti |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CPP-UT-Bench: Can LLMs Write Complex Unit Tests in C++?
by: Bhargava, Vaishnavi, et al.
Published: (2024)
by: Bhargava, Vaishnavi, et al.
Published: (2024)
BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
by: Zhou, Jinan, et al.
Published: (2025)
by: Zhou, Jinan, et al.
Published: (2025)
SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories
by: Soni, Aditya Bharat, et al.
Published: (2026)
by: Soni, Aditya Bharat, et al.
Published: (2026)
Action Shapley: A Training Data Selection Metric for World Model in Reinforcement Learning
by: Ghosh, Rajat, et al.
Published: (2026)
by: Ghosh, Rajat, et al.
Published: (2026)
Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go
by: Pipalani, Yashshi, et al.
Published: (2025)
by: Pipalani, Yashshi, et al.
Published: (2025)
Efficient Alignment of Large Language Models via Data Sampling
by: Khera, Amrit, et al.
Published: (2024)
by: Khera, Amrit, et al.
Published: (2024)
MLKV: Efficiently Scaling up Large Embedding Model Training with Disk-based Key-Value Storage
by: He, Yongjun, et al.
Published: (2025)
by: He, Yongjun, et al.
Published: (2025)
RANGER -- Repository-Level Agent for Graph-Enhanced Retrieval
by: Shah, Pratik, et al.
Published: (2025)
by: Shah, Pratik, et al.
Published: (2025)
A Multi-Agent Framework for Stateful Inference-Time Search
by: Lalan, Arshika, et al.
Published: (2025)
by: Lalan, Arshika, et al.
Published: (2025)
Transformation-Augmented GRPO for Enhancing Exploration in Reasoning of Large Language Models
by: Le, Khiem, et al.
Published: (2026)
by: Le, Khiem, et al.
Published: (2026)
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models
by: Ding, Fei, et al.
Published: (2025)
by: Ding, Fei, et al.
Published: (2025)
Scaling Law for Large-Scale Pre-Training Using Chaotic Time Series and Predictability in Financial Time Series
by: Takemoto, Yuki
Published: (2025)
by: Takemoto, Yuki
Published: (2025)
Scaling Laws for Post Training Quantized Large Language Models
by: Xu, Zifei, et al.
Published: (2024)
by: Xu, Zifei, et al.
Published: (2024)
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning
by: Chen, Xiwen, et al.
Published: (2025)
by: Chen, Xiwen, et al.
Published: (2025)
TEMPO: Scaling Test-time Training for Large Reasoning Models
by: Zhang, Qingyang, et al.
Published: (2026)
by: Zhang, Qingyang, et al.
Published: (2026)
Graph-GRPO: Training Graph Flow Models with Reinforcement Learning
by: Zhu, Baoheng, et al.
Published: (2026)
by: Zhu, Baoheng, et al.
Published: (2026)
S-GRPO: Unified Post-Training for Large Vision-Language Models
by: Yan, Yuming, et al.
Published: (2026)
by: Yan, Yuming, et al.
Published: (2026)
Exploring Scaling Laws for Local SGD in Large Language Model Training
by: He, Qiaozhi, et al.
Published: (2024)
by: He, Qiaozhi, et al.
Published: (2024)
Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward
by: Liu, Zikang, et al.
Published: (2025)
by: Liu, Zikang, et al.
Published: (2025)
WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning
by: Mundada, Gagan, et al.
Published: (2026)
by: Mundada, Gagan, et al.
Published: (2026)
A Hierarchical Language Model with Predictable Scaling Laws and Provable Benefits of Reasoning
by: Gaitonde, Jason, et al.
Published: (2026)
by: Gaitonde, Jason, et al.
Published: (2026)
ELLA: Efficient Lifelong Learning for Adapters in Large Language Models
by: Biswas, Shristi Das, et al.
Published: (2026)
by: Biswas, Shristi Das, et al.
Published: (2026)
BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
by: Li, Yuming, et al.
Published: (2025)
by: Li, Yuming, et al.
Published: (2025)
S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
by: Dai, Muzhi, et al.
Published: (2025)
by: Dai, Muzhi, et al.
Published: (2025)
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning
by: Wang, Hu, et al.
Published: (2025)
by: Wang, Hu, et al.
Published: (2025)
TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
by: Ding, Zheng, et al.
Published: (2025)
by: Ding, Zheng, et al.
Published: (2025)
Input Guided Multiple Deconstruction Single Reconstruction neural network models for Matrix Factorization
by: Dutta, Prasun, et al.
Published: (2024)
by: Dutta, Prasun, et al.
Published: (2024)
The ALCHEmist: Automated Labeling 500x CHEaper Than LLM Data Annotators
by: Huang, Tzu-Heng, et al.
Published: (2024)
by: Huang, Tzu-Heng, et al.
Published: (2024)
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
by: Qu, Yun, et al.
Published: (2026)
by: Qu, Yun, et al.
Published: (2026)
Forward-Cooperation-Backward (FCB) learning in a Multi-Encoding Uni-Decoding neural network architecture
by: Dutta, Prasun, et al.
Published: (2025)
by: Dutta, Prasun, et al.
Published: (2025)
FairGRPO: Fair Reinforcement Learning for Equitable Clinical Reasoning
by: Dai, Shiqi, et al.
Published: (2025)
by: Dai, Shiqi, et al.
Published: (2025)
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning
by: Tian, Minghao, et al.
Published: (2026)
by: Tian, Minghao, et al.
Published: (2026)
Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning
by: Wu, Xiaojun, et al.
Published: (2025)
by: Wu, Xiaojun, et al.
Published: (2025)
Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models
by: Costello, Caia, et al.
Published: (2025)
by: Costello, Caia, et al.
Published: (2025)
Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
by: Bereket, Michael, et al.
Published: (2025)
by: Bereket, Michael, et al.
Published: (2025)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
by: Parthasarathi, Prasanna, et al.
Published: (2025)
by: Parthasarathi, Prasanna, et al.
Published: (2025)
Training Large Reasoning Models Efficiently via Progressive Thought Encoding
by: Zhang, Zeliang, et al.
Published: (2026)
by: Zhang, Zeliang, et al.
Published: (2026)
Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
by: Hägele, Alexander, et al.
Published: (2024)
by: Hägele, Alexander, et al.
Published: (2024)
Scaling Law for Quantization-Aware Training
by: Chen, Mengzhao, et al.
Published: (2025)
by: Chen, Mengzhao, et al.
Published: (2025)
Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language Models
by: Lin, Junhong, et al.
Published: (2025)
by: Lin, Junhong, et al.
Published: (2025)
Similar Items
-
CPP-UT-Bench: Can LLMs Write Complex Unit Tests in C++?
by: Bhargava, Vaishnavi, et al.
Published: (2024) -
BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
by: Zhou, Jinan, et al.
Published: (2025) -
SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories
by: Soni, Aditya Bharat, et al.
Published: (2026) -
Action Shapley: A Training Data Selection Metric for World Model in Reinforcement Learning
by: Ghosh, Rajat, et al.
Published: (2026) -
Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go
by: Pipalani, Yashshi, et al.
Published: (2025)