Predictive Scaling Laws for Efficient GRPO Training of Large Reasoning Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Nimmaturi, Datta, Bhargava, Vaishnavi, Ghosh, Rajat, George, Johnu, Dutta, Debojyoti |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CPP-UT-Bench: Can LLMs Write Complex Unit Tests in C++?
di: Bhargava, Vaishnavi, et al.
Pubblicazione: (2024)
di: Bhargava, Vaishnavi, et al.
Pubblicazione: (2024)
BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
di: Zhou, Jinan, et al.
Pubblicazione: (2025)
di: Zhou, Jinan, et al.
Pubblicazione: (2025)
SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories
di: Soni, Aditya Bharat, et al.
Pubblicazione: (2026)
di: Soni, Aditya Bharat, et al.
Pubblicazione: (2026)
Action Shapley: A Training Data Selection Metric for World Model in Reinforcement Learning
di: Ghosh, Rajat, et al.
Pubblicazione: (2026)
di: Ghosh, Rajat, et al.
Pubblicazione: (2026)
Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go
di: Pipalani, Yashshi, et al.
Pubblicazione: (2025)
di: Pipalani, Yashshi, et al.
Pubblicazione: (2025)
Efficient Alignment of Large Language Models via Data Sampling
di: Khera, Amrit, et al.
Pubblicazione: (2024)
di: Khera, Amrit, et al.
Pubblicazione: (2024)
MLKV: Efficiently Scaling up Large Embedding Model Training with Disk-based Key-Value Storage
di: He, Yongjun, et al.
Pubblicazione: (2025)
di: He, Yongjun, et al.
Pubblicazione: (2025)
RANGER -- Repository-Level Agent for Graph-Enhanced Retrieval
di: Shah, Pratik, et al.
Pubblicazione: (2025)
di: Shah, Pratik, et al.
Pubblicazione: (2025)
A Multi-Agent Framework for Stateful Inference-Time Search
di: Lalan, Arshika, et al.
Pubblicazione: (2025)
di: Lalan, Arshika, et al.
Pubblicazione: (2025)
Transformation-Augmented GRPO for Enhancing Exploration in Reasoning of Large Language Models
di: Le, Khiem, et al.
Pubblicazione: (2026)
di: Le, Khiem, et al.
Pubblicazione: (2026)
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models
di: Ding, Fei, et al.
Pubblicazione: (2025)
di: Ding, Fei, et al.
Pubblicazione: (2025)
Scaling Law for Large-Scale Pre-Training Using Chaotic Time Series and Predictability in Financial Time Series
di: Takemoto, Yuki
Pubblicazione: (2025)
di: Takemoto, Yuki
Pubblicazione: (2025)
Scaling Laws for Post Training Quantized Large Language Models
di: Xu, Zifei, et al.
Pubblicazione: (2024)
di: Xu, Zifei, et al.
Pubblicazione: (2024)
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning
di: Chen, Xiwen, et al.
Pubblicazione: (2025)
di: Chen, Xiwen, et al.
Pubblicazione: (2025)
TEMPO: Scaling Test-time Training for Large Reasoning Models
di: Zhang, Qingyang, et al.
Pubblicazione: (2026)
di: Zhang, Qingyang, et al.
Pubblicazione: (2026)
Graph-GRPO: Training Graph Flow Models with Reinforcement Learning
di: Zhu, Baoheng, et al.
Pubblicazione: (2026)
di: Zhu, Baoheng, et al.
Pubblicazione: (2026)
S-GRPO: Unified Post-Training for Large Vision-Language Models
di: Yan, Yuming, et al.
Pubblicazione: (2026)
di: Yan, Yuming, et al.
Pubblicazione: (2026)
Exploring Scaling Laws for Local SGD in Large Language Model Training
di: He, Qiaozhi, et al.
Pubblicazione: (2024)
di: He, Qiaozhi, et al.
Pubblicazione: (2024)
Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward
di: Liu, Zikang, et al.
Pubblicazione: (2025)
di: Liu, Zikang, et al.
Pubblicazione: (2025)
WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning
di: Mundada, Gagan, et al.
Pubblicazione: (2026)
di: Mundada, Gagan, et al.
Pubblicazione: (2026)
A Hierarchical Language Model with Predictable Scaling Laws and Provable Benefits of Reasoning
di: Gaitonde, Jason, et al.
Pubblicazione: (2026)
di: Gaitonde, Jason, et al.
Pubblicazione: (2026)
ELLA: Efficient Lifelong Learning for Adapters in Large Language Models
di: Biswas, Shristi Das, et al.
Pubblicazione: (2026)
di: Biswas, Shristi Das, et al.
Pubblicazione: (2026)
BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
di: Li, Yuming, et al.
Pubblicazione: (2025)
di: Li, Yuming, et al.
Pubblicazione: (2025)
S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
di: Dai, Muzhi, et al.
Pubblicazione: (2025)
di: Dai, Muzhi, et al.
Pubblicazione: (2025)
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning
di: Wang, Hu, et al.
Pubblicazione: (2025)
di: Wang, Hu, et al.
Pubblicazione: (2025)
TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
di: Ding, Zheng, et al.
Pubblicazione: (2025)
di: Ding, Zheng, et al.
Pubblicazione: (2025)
Input Guided Multiple Deconstruction Single Reconstruction neural network models for Matrix Factorization
di: Dutta, Prasun, et al.
Pubblicazione: (2024)
di: Dutta, Prasun, et al.
Pubblicazione: (2024)
The ALCHEmist: Automated Labeling 500x CHEaper Than LLM Data Annotators
di: Huang, Tzu-Heng, et al.
Pubblicazione: (2024)
di: Huang, Tzu-Heng, et al.
Pubblicazione: (2024)
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
di: Qu, Yun, et al.
Pubblicazione: (2026)
di: Qu, Yun, et al.
Pubblicazione: (2026)
Forward-Cooperation-Backward (FCB) learning in a Multi-Encoding Uni-Decoding neural network architecture
di: Dutta, Prasun, et al.
Pubblicazione: (2025)
di: Dutta, Prasun, et al.
Pubblicazione: (2025)
FairGRPO: Fair Reinforcement Learning for Equitable Clinical Reasoning
di: Dai, Shiqi, et al.
Pubblicazione: (2025)
di: Dai, Shiqi, et al.
Pubblicazione: (2025)
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning
di: Tian, Minghao, et al.
Pubblicazione: (2026)
di: Tian, Minghao, et al.
Pubblicazione: (2026)
Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning
di: Wu, Xiaojun, et al.
Pubblicazione: (2025)
di: Wu, Xiaojun, et al.
Pubblicazione: (2025)
Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models
di: Costello, Caia, et al.
Pubblicazione: (2025)
di: Costello, Caia, et al.
Pubblicazione: (2025)
Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
di: Bereket, Michael, et al.
Pubblicazione: (2025)
di: Bereket, Michael, et al.
Pubblicazione: (2025)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
di: Parthasarathi, Prasanna, et al.
Pubblicazione: (2025)
di: Parthasarathi, Prasanna, et al.
Pubblicazione: (2025)
Training Large Reasoning Models Efficiently via Progressive Thought Encoding
di: Zhang, Zeliang, et al.
Pubblicazione: (2026)
di: Zhang, Zeliang, et al.
Pubblicazione: (2026)
Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
di: Hägele, Alexander, et al.
Pubblicazione: (2024)
di: Hägele, Alexander, et al.
Pubblicazione: (2024)
Scaling Law for Quantization-Aware Training
di: Chen, Mengzhao, et al.
Pubblicazione: (2025)
di: Chen, Mengzhao, et al.
Pubblicazione: (2025)
Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language Models
di: Lin, Junhong, et al.
Pubblicazione: (2025)
di: Lin, Junhong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CPP-UT-Bench: Can LLMs Write Complex Unit Tests in C++?
di: Bhargava, Vaishnavi, et al.
Pubblicazione: (2024) -
BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
di: Zhou, Jinan, et al.
Pubblicazione: (2025) -
SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories
di: Soni, Aditya Bharat, et al.
Pubblicazione: (2026) -
Action Shapley: A Training Data Selection Metric for World Model in Reinforcement Learning
di: Ghosh, Rajat, et al.
Pubblicazione: (2026) -
Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go
di: Pipalani, Yashshi, et al.
Pubblicazione: (2025)