Why GRPO Needs Normalization: A Local-Curvature Perspective on Adaptive Gradients
Fuente:
arXiv
Salvato in:
| Autori principali: | Ge, Cheng, Yin, Caitlyn Heqi, Liang, Hao, Zhang, Jiawei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Gradient Starvation in Binary-Reward GRPO: Why Group-Mean Centering Fails and Why the Simplest Fix Works
di: Nie, Wenhua, et al.
Pubblicazione: (2026)
di: Nie, Wenhua, et al.
Pubblicazione: (2026)
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning
di: Chen, Xiwen, et al.
Pubblicazione: (2025)
di: Chen, Xiwen, et al.
Pubblicazione: (2025)
Is Conformal Factuality for RAG-based LLMs Robust? Novel Metrics and Systematic Insights
di: Chen, Yi, et al.
Pubblicazione: (2026)
di: Chen, Yi, et al.
Pubblicazione: (2026)
Why Transformers Need Adam: A Hessian Perspective
di: Zhang, Yushun, et al.
Pubblicazione: (2024)
di: Zhang, Yushun, et al.
Pubblicazione: (2024)
Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty Adaptation
di: Yu, Zhiqi, et al.
Pubblicazione: (2026)
di: Yu, Zhiqi, et al.
Pubblicazione: (2026)
Why Do We Need Warm-up? A Theoretical Perspective
di: Alimisis, Foivos, et al.
Pubblicazione: (2025)
di: Alimisis, Foivos, et al.
Pubblicazione: (2025)
Adaptive Optimization via Momentum on Variance-Normalized Gradients
di: Patitucci, Francisco, et al.
Pubblicazione: (2026)
di: Patitucci, Francisco, et al.
Pubblicazione: (2026)
Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
di: Mansouri, Omar El, et al.
Pubblicazione: (2025)
di: Mansouri, Omar El, et al.
Pubblicazione: (2025)
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
di: Zhang, Zhenyu, et al.
Pubblicazione: (2024)
di: Zhang, Zhenyu, et al.
Pubblicazione: (2024)
LithoGRPO: Fast Inverse Lithography via GRPO Reinforced Flow Matching
di: Lai, Yao, et al.
Pubblicazione: (2026)
di: Lai, Yao, et al.
Pubblicazione: (2026)
A Locally Adaptive Algorithm for Multiple Testing with Network Structure
di: Liang, Ziyi, et al.
Pubblicazione: (2022)
di: Liang, Ziyi, et al.
Pubblicazione: (2022)
Stochastic Normalized Gradient Descent with Momentum for Large-Batch Training
di: Zhao, Shen-Yi, et al.
Pubblicazione: (2020)
di: Zhao, Shen-Yi, et al.
Pubblicazione: (2020)
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2026)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2026)
Minimax Optimality of Score-based Diffusion Models: Beyond the Density Lower Bound Assumptions
di: Zhang, Kaihong, et al.
Pubblicazione: (2024)
di: Zhang, Kaihong, et al.
Pubblicazione: (2024)
Reconstructing MODIS Normalized Difference Snow Index Product on Greenland Ice Sheet Using Spatiotemporal Extreme Gradient Boosting Model
di: Ye, Fan, et al.
Pubblicazione: (2024)
di: Ye, Fan, et al.
Pubblicazione: (2024)
Adaptive Gradient Normalization and Independent Sampling for (Stochastic) Generalized-Smooth Optimization
di: Yang, Yufeng, et al.
Pubblicazione: (2024)
di: Yang, Yufeng, et al.
Pubblicazione: (2024)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
di: Yari, Amir Hossein, et al.
Pubblicazione: (2026)
di: Yari, Amir Hossein, et al.
Pubblicazione: (2026)
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
di: Chen, Yi, et al.
Pubblicazione: (2025)
di: Chen, Yi, et al.
Pubblicazione: (2025)
Normal Guidance is what Attention Needs
di: Harvey, Ethan, et al.
Pubblicazione: (2026)
di: Harvey, Ethan, et al.
Pubblicazione: (2026)
Mastering AI: Big Data, Deep Learning, and the Evolution of Large Language Models -- AutoML from Basics to State-of-the-Art Techniques
di: Feng, Pohsun, et al.
Pubblicazione: (2024)
di: Feng, Pohsun, et al.
Pubblicazione: (2024)
FairGRPO: Fair Reinforcement Learning for Equitable Clinical Reasoning
di: Dai, Shiqi, et al.
Pubblicazione: (2025)
di: Dai, Shiqi, et al.
Pubblicazione: (2025)
AceGRPO: Adaptive Curriculum Enhanced Group Relative Policy Optimization for Autonomous Machine Learning Engineering
di: Cai, Yuzhu, et al.
Pubblicazione: (2026)
di: Cai, Yuzhu, et al.
Pubblicazione: (2026)
How Does Gradient Descent Learn Features -- A Local Analysis for Regularized Two-Layer Neural Networks
di: Zhou, Mo, et al.
Pubblicazione: (2024)
di: Zhou, Mo, et al.
Pubblicazione: (2024)
The Blessing of Dimensionality in LLM Fine-tuning: A Variance-Curvature Perspective
di: Liang, Qiyao, et al.
Pubblicazione: (2026)
di: Liang, Qiyao, et al.
Pubblicazione: (2026)
Hard Examples Are All You Need: Maximizing GRPO Post-Training Under Annotation Budgets
di: Pikus, Benjamin, et al.
Pubblicazione: (2025)
di: Pikus, Benjamin, et al.
Pubblicazione: (2025)
When and Why Does Unsupervised RL Succeed in Mathematical Reasoning? A Manifold Envelopment Perspective
di: Zhang, Zelin, et al.
Pubblicazione: (2026)
di: Zhang, Zelin, et al.
Pubblicazione: (2026)
Unsupervised Adaptive Normalization
di: Faye, Bilal, et al.
Pubblicazione: (2024)
di: Faye, Bilal, et al.
Pubblicazione: (2024)
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
di: Jin, Ruinan, et al.
Pubblicazione: (2026)
di: Jin, Ruinan, et al.
Pubblicazione: (2026)
Advances in GRPO for Generation Models: A Survey
di: Liu, Zexiang, et al.
Pubblicazione: (2026)
di: Liu, Zexiang, et al.
Pubblicazione: (2026)
MAC: An Efficient Gradient Preconditioning using Mean Activation Approximated Curvature
di: Seung, Hyunseok, et al.
Pubblicazione: (2025)
di: Seung, Hyunseok, et al.
Pubblicazione: (2025)
CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
di: Tabesh, Soroush, et al.
Pubblicazione: (2025)
di: Tabesh, Soroush, et al.
Pubblicazione: (2025)
A Lightweight and Gradient-Stable Neural Layer
di: Yu, Yueyao, et al.
Pubblicazione: (2021)
di: Yu, Yueyao, et al.
Pubblicazione: (2021)
F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking
di: Surana, Rohan, et al.
Pubblicazione: (2026)
di: Surana, Rohan, et al.
Pubblicazione: (2026)
Exploiting Curvature in Online Convex Optimization with Delayed Feedback
di: Qiu, Hao, et al.
Pubblicazione: (2025)
di: Qiu, Hao, et al.
Pubblicazione: (2025)
GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
di: Xu, Yanchen, et al.
Pubblicazione: (2025)
di: Xu, Yanchen, et al.
Pubblicazione: (2025)
Position: Why a Dynamical Systems Perspective is Needed to Advance Time Series Modeling
di: Durstewitz, Daniel, et al.
Pubblicazione: (2026)
di: Durstewitz, Daniel, et al.
Pubblicazione: (2026)
CAT: Curvature-Adaptive Transformers for Geometry-Aware Learning
di: Lin, Ryan Y., et al.
Pubblicazione: (2025)
di: Lin, Ryan Y., et al.
Pubblicazione: (2025)
Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent
di: Zhang, Chenyang, et al.
Pubblicazione: (2026)
di: Zhang, Chenyang, et al.
Pubblicazione: (2026)
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning
di: Tian, Minghao, et al.
Pubblicazione: (2026)
di: Tian, Minghao, et al.
Pubblicazione: (2026)
Graph-GRPO: Training Graph Flow Models with Reinforcement Learning
di: Zhu, Baoheng, et al.
Pubblicazione: (2026)
di: Zhu, Baoheng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Gradient Starvation in Binary-Reward GRPO: Why Group-Mean Centering Fails and Why the Simplest Fix Works
di: Nie, Wenhua, et al.
Pubblicazione: (2026) -
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning
di: Chen, Xiwen, et al.
Pubblicazione: (2025) -
Is Conformal Factuality for RAG-based LLMs Robust? Novel Metrics and Systematic Insights
di: Chen, Yi, et al.
Pubblicazione: (2026) -
Why Transformers Need Adam: A Hessian Perspective
di: Zhang, Yushun, et al.
Pubblicazione: (2024) -
Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty Adaptation
di: Yu, Zhiqi, et al.
Pubblicazione: (2026)