Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mansouri, Omar El, Izzati, Fathinah Asma, Seddik, Mohamed El Amine, Lahlou, Salem |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
High-dimensional Learning with Noisy Labels
von: Firdoussi, Aymane El, et al.
Veröffentlicht: (2024)
von: Firdoussi, Aymane El, et al.
Veröffentlicht: (2024)
$α$-LoRA: Effective Fine-Tuning via Base Model Rescaling
von: Firdoussi, Aymane El, et al.
Veröffentlicht: (2025)
von: Firdoussi, Aymane El, et al.
Veröffentlicht: (2025)
Investigating Regularization of Self-Play Language Models
von: Alami, Reda, et al.
Veröffentlicht: (2024)
von: Alami, Reda, et al.
Veröffentlicht: (2024)
Maximizing the Potential of Synthetic Data: Insights from Random Matrix Theory
von: Firdoussi, Aymane El, et al.
Veröffentlicht: (2024)
von: Firdoussi, Aymane El, et al.
Veröffentlicht: (2024)
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
von: Seddik, Mohamed El Amine, et al.
Veröffentlicht: (2024)
von: Seddik, Mohamed El Amine, et al.
Veröffentlicht: (2024)
How Does Attention Help? Insights from Random Matrices on Signal Recovery from Sequence Models
von: Seddik, Mohamed El Amine
Veröffentlicht: (2026)
von: Seddik, Mohamed El Amine
Veröffentlicht: (2026)
Accurate and Diverse LLM Mathematical Reasoning via Automated PRM-Guided GFlowNets
von: Younsi, Adam, et al.
Veröffentlicht: (2025)
von: Younsi, Adam, et al.
Veröffentlicht: (2025)
On the Privacy Risks of Spiking Neural Networks: A Membership Inference Analysis
von: Guan, Junyi, et al.
Veröffentlicht: (2025)
von: Guan, Junyi, et al.
Veröffentlicht: (2025)
GRPO is Secretly a Process Reward Model
von: Sullivan, Michael, et al.
Veröffentlicht: (2025)
von: Sullivan, Michael, et al.
Veröffentlicht: (2025)
Improved Exploration in GFlownets via Enhanced Epistemic Neural Networks
von: Muhammad, Sajan, et al.
Veröffentlicht: (2025)
von: Muhammad, Sajan, et al.
Veröffentlicht: (2025)
Stabilizing Policy Gradient Methods via Reward Profiling
von: Ahmed, Shihab, et al.
Veröffentlicht: (2025)
von: Ahmed, Shihab, et al.
Veröffentlicht: (2025)
From Reward Shaping to Q-Shaping: Achieving Unbiased Learning with LLM-Guided Knowledge
von: Wu, Xiefeng
Veröffentlicht: (2024)
von: Wu, Xiefeng
Veröffentlicht: (2024)
GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification
von: Gan, Wangjie, et al.
Veröffentlicht: (2026)
von: Gan, Wangjie, et al.
Veröffentlicht: (2026)
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation
von: Izzati, Fathinah, et al.
Veröffentlicht: (2025)
von: Izzati, Fathinah, et al.
Veröffentlicht: (2025)
Fréchet regression with implicit denoising and multicollinearity reduction
von: Mansouri, Dou El Kefel, et al.
Veröffentlicht: (2024)
von: Mansouri, Dou El Kefel, et al.
Veröffentlicht: (2024)
Quantum Agents for Algorithmic Discovery
von: Kerenidis, Iordanis, et al.
Veröffentlicht: (2025)
von: Kerenidis, Iordanis, et al.
Veröffentlicht: (2025)
Zero-Shot Off-Policy Learning
von: Asadulaev, Arip, et al.
Veröffentlicht: (2026)
von: Asadulaev, Arip, et al.
Veröffentlicht: (2026)
MusiScene: Leveraging MU-LLaMA for Scene Imagination and Enhanced Video Background Music Generation
von: Izzati, Fathinah, et al.
Veröffentlicht: (2025)
von: Izzati, Fathinah, et al.
Veröffentlicht: (2025)
Unbiased Gradient Low-Rank Projection
von: Pan, Rui, et al.
Veröffentlicht: (2025)
von: Pan, Rui, et al.
Veröffentlicht: (2025)
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
von: Chmiel, Brian, et al.
Veröffentlicht: (2022)
von: Chmiel, Brian, et al.
Veröffentlicht: (2022)
GFlowNet Foundations
von: Bengio, Yoshua, et al.
Veröffentlicht: (2021)
von: Bengio, Yoshua, et al.
Veröffentlicht: (2021)
MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
von: Wei, Kangda, et al.
Veröffentlicht: (2026)
von: Wei, Kangda, et al.
Veröffentlicht: (2026)
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2025)
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2025)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2026)
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2026)
VARS-FL: Validation-Aligned Client Selection for Non-IID Federated Learning in IoT Systems
von: Lakas, Mohamed, et al.
Veröffentlicht: (2026)
von: Lakas, Mohamed, et al.
Veröffentlicht: (2026)
An Imperfect Verifier is Good Enough: Learning with Noisy Rewards
von: Plesner, Andreas, et al.
Veröffentlicht: (2026)
von: Plesner, Andreas, et al.
Veröffentlicht: (2026)
Learning Robust Reward Machines from Noisy Labels
von: Parac, Roko, et al.
Veröffentlicht: (2024)
von: Parac, Roko, et al.
Veröffentlicht: (2024)
Epileptic Seizure Prediction Using Patient-Adaptive Transformer Networks
von: Mahdi, Mohamed, et al.
Veröffentlicht: (2026)
von: Mahdi, Mohamed, et al.
Veröffentlicht: (2026)
Reward Learning from Multiple Feedback Types
von: Metz, Yannick, et al.
Veröffentlicht: (2025)
von: Metz, Yannick, et al.
Veröffentlicht: (2025)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
Panther: Faster and Cheaper Computations with Randomized Numerical Linear Algebra
von: Seddik, Fahd, et al.
Veröffentlicht: (2026)
von: Seddik, Fahd, et al.
Veröffentlicht: (2026)
Optimal Transport for LLM Reward Modeling from Noisy Preference
von: Pan, Licheng, et al.
Veröffentlicht: (2026)
von: Pan, Licheng, et al.
Veröffentlicht: (2026)
Mitigating Societal Cognitive Overload in the Age of AI: Challenges and Directions
von: Lahlou, Salem
Veröffentlicht: (2025)
von: Lahlou, Salem
Veröffentlicht: (2025)
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
von: Cai, Xin-Qiang, et al.
Veröffentlicht: (2025)
von: Cai, Xin-Qiang, et al.
Veröffentlicht: (2025)
Gradient-Free Noise Optimization for Reward Alignment in Generative Models
von: Kim, Jeongsol, et al.
Veröffentlicht: (2026)
von: Kim, Jeongsol, et al.
Veröffentlicht: (2026)
On the Optimal Construction of Unbiased Gradient Estimators for Zeroth-Order Optimization
von: Ma, Shaocong, et al.
Veröffentlicht: (2025)
von: Ma, Shaocong, et al.
Veröffentlicht: (2025)
Do Vision and Language Encoders Represent the World Similarly?
von: Maniparambil, Mayug, et al.
Veröffentlicht: (2024)
von: Maniparambil, Mayug, et al.
Veröffentlicht: (2024)
What is the Alignment Objective of GRPO?
von: Vojnovic, Milan, et al.
Veröffentlicht: (2025)
von: Vojnovic, Milan, et al.
Veröffentlicht: (2025)
Learning with Noisy Labels by Adaptive Gradient-Based Outlier Removal
von: Sedova, Anastasiia, et al.
Veröffentlicht: (2023)
von: Sedova, Anastasiia, et al.
Veröffentlicht: (2023)
MixtureKit: A General Framework for Composing, Training, and Visualizing Mixture-of-Experts Models
von: Chamma, Ahmad, et al.
Veröffentlicht: (2025)
von: Chamma, Ahmad, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
High-dimensional Learning with Noisy Labels
von: Firdoussi, Aymane El, et al.
Veröffentlicht: (2024) -
$α$-LoRA: Effective Fine-Tuning via Base Model Rescaling
von: Firdoussi, Aymane El, et al.
Veröffentlicht: (2025) -
Investigating Regularization of Self-Play Language Models
von: Alami, Reda, et al.
Veröffentlicht: (2024) -
Maximizing the Potential of Synthetic Data: Insights from Random Matrix Theory
von: Firdoussi, Aymane El, et al.
Veröffentlicht: (2024) -
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
von: Seddik, Mohamed El Amine, et al.
Veröffentlicht: (2024)