Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hoy, William, Wang, Binxu, Pan, Xu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LithoGRPO: Fast Inverse Lithography via GRPO Reinforced Flow Matching
von: Lai, Yao, et al.
Veröffentlicht: (2026)
von: Lai, Yao, et al.
Veröffentlicht: (2026)
S-GRPO: Unified Post-Training for Large Vision-Language Models
von: Yan, Yuming, et al.
Veröffentlicht: (2026)
von: Yan, Yuming, et al.
Veröffentlicht: (2026)
TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
von: Ding, Zheng, et al.
Veröffentlicht: (2025)
von: Ding, Zheng, et al.
Veröffentlicht: (2025)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training
von: Wang, Yuanyi, et al.
Veröffentlicht: (2026)
von: Wang, Yuanyi, et al.
Veröffentlicht: (2026)
SFT-GRPO Data Overlap as a Post-Training Hyperparameter for Autoformalization
von: Su, Xiaole, et al.
Veröffentlicht: (2026)
von: Su, Xiaole, et al.
Veröffentlicht: (2026)
Graph-GRPO: Training Graph Flow Models with Reinforcement Learning
von: Zhu, Baoheng, et al.
Veröffentlicht: (2026)
von: Zhu, Baoheng, et al.
Veröffentlicht: (2026)
Hard Examples Are All You Need: Maximizing GRPO Post-Training Under Annotation Budgets
von: Pikus, Benjamin, et al.
Veröffentlicht: (2025)
von: Pikus, Benjamin, et al.
Veröffentlicht: (2025)
Circuit Mechanisms for Spatial Relation Generation in Diffusion Transformers
von: Wang, Binxu, et al.
Veröffentlicht: (2026)
von: Wang, Binxu, et al.
Veröffentlicht: (2026)
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning
von: Tian, Minghao, et al.
Veröffentlicht: (2026)
von: Tian, Minghao, et al.
Veröffentlicht: (2026)
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training
von: Hu, Zhengding, et al.
Veröffentlicht: (2026)
von: Hu, Zhengding, et al.
Veröffentlicht: (2026)
MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
Stepwise Credit Assignment for GRPO on Flow-Matching Models
von: Savani, Yash, et al.
Veröffentlicht: (2026)
von: Savani, Yash, et al.
Veröffentlicht: (2026)
STABLE: Gated Continual Learning for Large Language Models
von: Hoy, William, et al.
Veröffentlicht: (2025)
von: Hoy, William, et al.
Veröffentlicht: (2025)
Predictive Scaling Laws for Efficient GRPO Training of Large Reasoning Models
von: Nimmaturi, Datta, et al.
Veröffentlicht: (2025)
von: Nimmaturi, Datta, et al.
Veröffentlicht: (2025)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
von: Parthasarathi, Prasanna, et al.
Veröffentlicht: (2025)
von: Parthasarathi, Prasanna, et al.
Veröffentlicht: (2025)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
von: Rank, Ben, et al.
Veröffentlicht: (2026)
von: Rank, Ben, et al.
Veröffentlicht: (2026)
Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
von: Hu, Pingbang, et al.
Veröffentlicht: (2026)
von: Hu, Pingbang, et al.
Veröffentlicht: (2026)
Rethinking Local Learning: A Cheaper and Faster Recipe for LLM Post-Training
von: Shi, Hengyu, et al.
Veröffentlicht: (2026)
von: Shi, Hengyu, et al.
Veröffentlicht: (2026)
MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
von: Wei, Kangda, et al.
Veröffentlicht: (2026)
von: Wei, Kangda, et al.
Veröffentlicht: (2026)
Accuracy vs. Accuracy: Computational Tradeoffs Between Classification Rates and Utility
von: Amit, Noga, et al.
Veröffentlicht: (2025)
von: Amit, Noga, et al.
Veröffentlicht: (2025)
GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping
von: Wang, Jing, et al.
Veröffentlicht: (2025)
von: Wang, Jing, et al.
Veröffentlicht: (2025)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2026)
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2026)
Revisiting MoE and Dense Speed-Accuracy Comparisons for LLM Training
von: Du, Xianzhi, et al.
Veröffentlicht: (2024)
von: Du, Xianzhi, et al.
Veröffentlicht: (2024)
On the Evolution of Federated Post-Training Large Language Models: A Model Accessibility View
von: Guo, Tao, et al.
Veröffentlicht: (2025)
von: Guo, Tao, et al.
Veröffentlicht: (2025)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning
von: Chen, Xiwen, et al.
Veröffentlicht: (2025)
von: Chen, Xiwen, et al.
Veröffentlicht: (2025)
Prompt Curriculum Learning for Efficient LLM Post-Training
von: Gao, Zhaolin, et al.
Veröffentlicht: (2025)
von: Gao, Zhaolin, et al.
Veröffentlicht: (2025)
Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward
von: Liu, Zikang, et al.
Veröffentlicht: (2025)
von: Liu, Zikang, et al.
Veröffentlicht: (2025)
f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment
von: Haldar, Rajdeep, et al.
Veröffentlicht: (2026)
von: Haldar, Rajdeep, et al.
Veröffentlicht: (2026)
CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training
von: Thede, Lukas, et al.
Veröffentlicht: (2026)
von: Thede, Lukas, et al.
Veröffentlicht: (2026)
Consolidating Rewarded Perturbations for LLM Post-Training
von: Zhang, Zheyu, et al.
Veröffentlicht: (2026)
von: Zhang, Zheyu, et al.
Veröffentlicht: (2026)
Automatic Configuration of LLM Post-Training Pipelines
von: Chwa, Channe, et al.
Veröffentlicht: (2026)
von: Chwa, Channe, et al.
Veröffentlicht: (2026)
Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
von: Rajani, Neel, et al.
Veröffentlicht: (2025)
von: Rajani, Neel, et al.
Veröffentlicht: (2025)
GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
von: Xu, Yanchen, et al.
Veröffentlicht: (2025)
von: Xu, Yanchen, et al.
Veröffentlicht: (2025)
AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
von: Han, Zhenyu, et al.
Veröffentlicht: (2025)
von: Han, Zhenyu, et al.
Veröffentlicht: (2025)
Approximation of Log-Partition Function in Policy Mirror Descent Induces Implicit Regularization for LLM Post-Training
von: Xu, Zhenghao, et al.
Veröffentlicht: (2026)
von: Xu, Zhenghao, et al.
Veröffentlicht: (2026)
Reinforce Adjoint Matching: Scaling RL Post-Training of Diffusion and Flow-Matching Models
von: Bergmeister, Andreas, et al.
Veröffentlicht: (2026)
von: Bergmeister, Andreas, et al.
Veröffentlicht: (2026)
Comparative Analysis and Parametric Tuning of PPO, GRPO, and DAPO for LLM Reasoning Enhancement
von: Lian, Yongsheng
Veröffentlicht: (2025)
von: Lian, Yongsheng
Veröffentlicht: (2025)
Elucidating Flow Matching ODE Dynamics with Respect to Data Geometries and Denoisers
von: Wan, Zhengchao, et al.
Veröffentlicht: (2024)
von: Wan, Zhengchao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LithoGRPO: Fast Inverse Lithography via GRPO Reinforced Flow Matching
von: Lai, Yao, et al.
Veröffentlicht: (2026) -
S-GRPO: Unified Post-Training for Large Vision-Language Models
von: Yan, Yuming, et al.
Veröffentlicht: (2026) -
TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
von: Ding, Zheng, et al.
Veröffentlicht: (2025) -
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
von: Xu, Yuanda, et al.
Veröffentlicht: (2026) -
Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training
von: Wang, Yuanyi, et al.
Veröffentlicht: (2026)