Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mroueh, Youssef, Dupuis, Nicolas, Belgodere, Brian, Nitsure, Apoorva, Rigotti, Mattia, Greenewald, Kristjan, Navratil, Jiri, Ross, Jerret, Rios, Jesus |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Risk Aware Benchmarking of Large Language Models
von: Nitsure, Apoorva, et al.
Veröffentlicht: (2023)
von: Nitsure, Apoorva, et al.
Veröffentlicht: (2023)
Distributional Preference Alignment of LLMs via Optimal Transport
von: Melnyk, Igor, et al.
Veröffentlicht: (2024)
von: Melnyk, Igor, et al.
Veröffentlicht: (2024)
Multivariate Stochastic Dominance via Optimal Transport and Applications to Models Benchmarking
von: Rioux, Gabriel, et al.
Veröffentlicht: (2024)
von: Rioux, Gabriel, et al.
Veröffentlicht: (2024)
GP-MoLFormer: A Foundation Model For Molecular Generation
von: Ross, Jerret, et al.
Veröffentlicht: (2024)
von: Ross, Jerret, et al.
Veröffentlicht: (2024)
Auditing and Generating Synthetic Data with Controllable Trust Trade-offs
von: Belgodere, Brian, et al.
Veröffentlicht: (2023)
von: Belgodere, Brian, et al.
Veröffentlicht: (2023)
GP-MoLFormer-Sim: Test Time Molecular Optimization through Contextual Similarity Guidance
von: Navratil, Jiri, et al.
Veröffentlicht: (2025)
von: Navratil, Jiri, et al.
Veröffentlicht: (2025)
Verify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
Information Theoretic Guarantees For Policy Alignment In Large Language Models
von: Mroueh, Youssef
Veröffentlicht: (2024)
von: Mroueh, Youssef
Veröffentlicht: (2024)
CliffSearch: Structured Agentic Co-Evolution over Theory and Code for Scientific Algorithm Discovery
von: Mroueh, Youssef, et al.
Veröffentlicht: (2026)
von: Mroueh, Youssef, et al.
Veröffentlicht: (2026)
Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge
von: Dognin, Pierre, et al.
Veröffentlicht: (2020)
von: Dognin, Pierre, et al.
Veröffentlicht: (2020)
Gradient Flows and Riemannian Structure in the Gromov-Wasserstein Geometry
von: Zhang, Zhengxin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhengxin, et al.
Veröffentlicht: (2024)
Quantum Verifiable Rewards for Post-Training Qiskit Code Assistant
von: Dupuis, Nicolas, et al.
Veröffentlicht: (2025)
von: Dupuis, Nicolas, et al.
Veröffentlicht: (2025)
Finite sample rates of convergence for the Bigraphical and Tensor graphical Lasso estimators
von: Zhou, Shuheng, et al.
Veröffentlicht: (2023)
von: Zhou, Shuheng, et al.
Veröffentlicht: (2023)
Reinforcement Learning with Verifiable Rewards: GRPO's Effective Loss, Dynamics, and Success Amplification
von: Mroueh, Youssef
Veröffentlicht: (2025)
von: Mroueh, Youssef
Veröffentlicht: (2025)
Training-Free Group Relative Policy Optimization
von: Cai, Yuzheng, et al.
Veröffentlicht: (2025)
von: Cai, Yuzheng, et al.
Veröffentlicht: (2025)
Privacy without Noisy Gradients: Slicing Mechanism for Generative Model Training
von: Greenewald, Kristjan, et al.
Veröffentlicht: (2024)
von: Greenewald, Kristjan, et al.
Veröffentlicht: (2024)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
von: Fakoor, Rasool, et al.
Veröffentlicht: (2026)
von: Fakoor, Rasool, et al.
Veröffentlicht: (2026)
Private Continuous-Time Synthetic Trajectory Generation via Mean-Field Langevin Dynamics
von: Gu, Anming, et al.
Veröffentlicht: (2025)
von: Gu, Anming, et al.
Veröffentlicht: (2025)
Partially Observed Trajectory Inference using Optimal Transport and a Dynamics Prior
von: Gu, Anming, et al.
Veröffentlicht: (2024)
von: Gu, Anming, et al.
Veröffentlicht: (2024)
Synthetic Census Data Generation via Multidimensional Multiset Sum
von: Dwork, Cynthia, et al.
Veröffentlicht: (2024)
von: Dwork, Cynthia, et al.
Veröffentlicht: (2024)
Domain Adaptable Prescriptive AI Agent for Enterprise
von: Orderique, Piero, et al.
Veröffentlicht: (2024)
von: Orderique, Piero, et al.
Veröffentlicht: (2024)
Logging Policy Design for Off-Policy Evaluation
von: Douglas, Connor, et al.
Veröffentlicht: (2026)
von: Douglas, Connor, et al.
Veröffentlicht: (2026)
Off-Policy Evaluation for Ranking Policies under Deterministic Logging Policies
von: Tanaka, Koichi, et al.
Veröffentlicht: (2026)
von: Tanaka, Koichi, et al.
Veröffentlicht: (2026)
When Do Off-Policy and On-Policy Policy Gradient Methods Align?
von: Mambelli, Davide, et al.
Veröffentlicht: (2024)
von: Mambelli, Davide, et al.
Veröffentlicht: (2024)
The Impact of Off-Policy Training Data on Probe Generalisation
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning
von: Zhao, Shiwan, et al.
Veröffentlicht: (2026)
von: Zhao, Shiwan, et al.
Veröffentlicht: (2026)
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
Constrained Group Relative Policy Optimization
von: Girgis, Roger, et al.
Veröffentlicht: (2026)
von: Girgis, Roger, et al.
Veröffentlicht: (2026)
Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
von: Ma, Rachel, et al.
Veröffentlicht: (2026)
von: Ma, Rachel, et al.
Veröffentlicht: (2026)
Fairness Is More Than Algorithms: Racial Disparities in Time-to-Recidivism
von: Han, Jessy Xinyi, et al.
Veröffentlicht: (2025)
von: Han, Jessy Xinyi, et al.
Veröffentlicht: (2025)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
von: Meng, Wenjia, et al.
Veröffentlicht: (2024)
von: Meng, Wenjia, et al.
Veröffentlicht: (2024)
Group-in-Group Policy Optimization for LLM Agent Training
von: Feng, Lang, et al.
Veröffentlicht: (2025)
von: Feng, Lang, et al.
Veröffentlicht: (2025)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
von: Lobo, Elita, et al.
Veröffentlicht: (2024)
von: Lobo, Elita, et al.
Veröffentlicht: (2024)
Residual Off-Policy RL for Finetuning Behavior Cloning Policies
von: Ankile, Lars, et al.
Veröffentlicht: (2025)
von: Ankile, Lars, et al.
Veröffentlicht: (2025)
Doubly-Robust Off-Policy Evaluation with Estimated Logging Policy
von: Lee, Kyungbok, et al.
Veröffentlicht: (2024)
von: Lee, Kyungbok, et al.
Veröffentlicht: (2024)
CPPO: Accelerating the Training of Group Relative Policy Optimization-Based Reasoning Models
von: Lin, Zhihang, et al.
Veröffentlicht: (2025)
von: Lin, Zhihang, et al.
Veröffentlicht: (2025)
Meta Off-Policy Estimation
von: Jeunen, Olivier
Veröffentlicht: (2025)
von: Jeunen, Olivier
Veröffentlicht: (2025)
Group Relative Policy Optimization for Image Captioning
von: Liang, Xu
Veröffentlicht: (2025)
von: Liang, Xu
Veröffentlicht: (2025)
Group Relative Policy Optimization for Speech Recognition
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2025)
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2025)
RePO: Bridging On-Policy Learning and Off-Policy Knowledge through Rephrasing Policy Optimization
von: Xia, Linxuan, et al.
Veröffentlicht: (2026)
von: Xia, Linxuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Risk Aware Benchmarking of Large Language Models
von: Nitsure, Apoorva, et al.
Veröffentlicht: (2023) -
Distributional Preference Alignment of LLMs via Optimal Transport
von: Melnyk, Igor, et al.
Veröffentlicht: (2024) -
Multivariate Stochastic Dominance via Optimal Transport and Applications to Models Benchmarking
von: Rioux, Gabriel, et al.
Veröffentlicht: (2024) -
GP-MoLFormer: A Foundation Model For Molecular Generation
von: Ross, Jerret, et al.
Veröffentlicht: (2024) -
Auditing and Generating Synthetic Data with Controllable Trust Trade-offs
von: Belgodere, Brian, et al.
Veröffentlicht: (2023)