GOPO: Policy Optimization using Ranked Rewards
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Choi, Kyuseong, Saha, Dwaipayan, Kim, Woojeong, Agarwal, Anish, Dwivedi, Raaz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TabImpute: Universal Zero-Shot Imputation for Tabular Data
von: Feitelberg, Jacob, et al.
Veröffentlicht: (2025)
von: Feitelberg, Jacob, et al.
Veröffentlicht: (2025)
Distributional Matrix Completion via Nearest Neighbors in the Wasserstein Space
von: Feitelberg, Jacob, et al.
Veröffentlicht: (2024)
von: Feitelberg, Jacob, et al.
Veröffentlicht: (2024)
Learning Counterfactual Distributions via Kernel Nearest Neighbors
von: Choi, Kyuseong, et al.
Veröffentlicht: (2024)
von: Choi, Kyuseong, et al.
Veröffentlicht: (2024)
Supervised Kernel Thinning
von: Gong, Albert, et al.
Veröffentlicht: (2024)
von: Gong, Albert, et al.
Veröffentlicht: (2024)
Doubly Robust Inference in Causal Latent Factor Models
von: Abadie, Alberto, et al.
Veröffentlicht: (2024)
von: Abadie, Alberto, et al.
Veröffentlicht: (2024)
N$^2$: A Unified Python Package and Test Bench for Nearest Neighbor-Based Matrix Completion
von: Chin, Caleb, et al.
Veröffentlicht: (2025)
von: Chin, Caleb, et al.
Veröffentlicht: (2025)
Intrinsic Reward Policy Optimization for Sparse-Reward Environments
von: Cho, Minjae, et al.
Veröffentlicht: (2026)
von: Cho, Minjae, et al.
Veröffentlicht: (2026)
From Chat Logs to Collective Insights: Aggregative Question Answering
von: Zhang, Wentao, et al.
Veröffentlicht: (2025)
von: Zhang, Wentao, et al.
Veröffentlicht: (2025)
An Imperfect Verifier is Good Enough: Learning with Noisy Rewards
von: Plesner, Andreas, et al.
Veröffentlicht: (2026)
von: Plesner, Andreas, et al.
Veröffentlicht: (2026)
Fairness Aware Reward Optimization
von: Choi, Ching Lam, et al.
Veröffentlicht: (2026)
von: Choi, Ching Lam, et al.
Veröffentlicht: (2026)
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
von: Huang, Chenghua, et al.
Veröffentlicht: (2025)
von: Huang, Chenghua, et al.
Veröffentlicht: (2025)
Value-Free Policy Optimization via Reward Partitioning
von: Faye, Bilal, et al.
Veröffentlicht: (2025)
von: Faye, Bilal, et al.
Veröffentlicht: (2025)
Mitigating Preference Hacking in Policy Optimization with Pessimism
von: Gupta, Dhawal, et al.
Veröffentlicht: (2025)
von: Gupta, Dhawal, et al.
Veröffentlicht: (2025)
ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization
von: Patel, Nirmal, et al.
Veröffentlicht: (2026)
von: Patel, Nirmal, et al.
Veröffentlicht: (2026)
CROP: Conservative Reward for Model-based Offline Policy Optimization
von: Li, Hao, et al.
Veröffentlicht: (2023)
von: Li, Hao, et al.
Veröffentlicht: (2023)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
von: Li, Gang, et al.
Veröffentlicht: (2025)
von: Li, Gang, et al.
Veröffentlicht: (2025)
Matrix Low-Rank Trust Region Policy Optimization
von: Rozada, Sergio, et al.
Veröffentlicht: (2024)
von: Rozada, Sergio, et al.
Veröffentlicht: (2024)
ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
von: Zhang, Chen Bo Calvin, et al.
Veröffentlicht: (2024)
von: Zhang, Chen Bo Calvin, et al.
Veröffentlicht: (2024)
Overcoming Reward Overoptimization via Adversarial Policy Optimization with Lightweight Uncertainty Estimation
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
ECPO: Evidence-Coupled Policy Optimization for Evidence-Certified Candidate Ranking
von: Hu, Miaobo, et al.
Veröffentlicht: (2026)
von: Hu, Miaobo, et al.
Veröffentlicht: (2026)
Gradient Extrapolation-Based Policy Optimization
von: Swapnil, Ismam Nur, et al.
Veröffentlicht: (2026)
von: Swapnil, Ismam Nur, et al.
Veröffentlicht: (2026)
Reward Learning through Ranking Mean Squared Error
von: Kharyal, Chaitanya, et al.
Veröffentlicht: (2026)
von: Kharyal, Chaitanya, et al.
Veröffentlicht: (2026)
Reliable Policy Iteration: Performance Robustness Across Architecture and Environment Perturbations
von: Eshwar, S. R., et al.
Veröffentlicht: (2025)
von: Eshwar, S. R., et al.
Veröffentlicht: (2025)
Pessimistic Off-Policy Optimization for Learning to Rank
von: Cief, Matej, et al.
Veröffentlicht: (2022)
von: Cief, Matej, et al.
Veröffentlicht: (2022)
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime
von: Sana, Mohamed, et al.
Veröffentlicht: (2026)
von: Sana, Mohamed, et al.
Veröffentlicht: (2026)
Trust Region Reward Optimization and Proximal Inverse Reward Optimization Algorithm
von: Chen, Yang, et al.
Veröffentlicht: (2025)
von: Chen, Yang, et al.
Veröffentlicht: (2025)
Clustered Policy Decision Ranking
von: Levin, Mark, et al.
Veröffentlicht: (2023)
von: Levin, Mark, et al.
Veröffentlicht: (2023)
RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training
von: Ren, Tao, et al.
Veröffentlicht: (2025)
von: Ren, Tao, et al.
Veröffentlicht: (2025)
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards
von: Ahmad, Ahmad, et al.
Veröffentlicht: (2024)
von: Ahmad, Ahmad, et al.
Veröffentlicht: (2024)
ReDit: Reward Dithering for Improved LLM Policy Optimization
von: Wei, Chenxing, et al.
Veröffentlicht: (2025)
von: Wei, Chenxing, et al.
Veröffentlicht: (2025)
WARP: On the Benefits of Weight Averaged Rewarded Policies
von: Ramé, Alexandre, et al.
Veröffentlicht: (2024)
von: Ramé, Alexandre, et al.
Veröffentlicht: (2024)
R^3: Replay, Reflection, and Ranking Rewards for LLM Reinforcement Learning
von: Jiang, Zhizheng, et al.
Veröffentlicht: (2026)
von: Jiang, Zhizheng, et al.
Veröffentlicht: (2026)
Optimizing Classification of Infrequent Labels by Reducing Variability in Label Distribution
von: Agarwal, Ashutosh
Veröffentlicht: (2025)
von: Agarwal, Ashutosh
Veröffentlicht: (2025)
Joint Optimization of Multi-Objective Reinforcement Learning with Policy Gradient Based Algorithm
von: Bai, Qinbo, et al.
Veröffentlicht: (2021)
von: Bai, Qinbo, et al.
Veröffentlicht: (2021)
Reward Hacking Mitigation using Verifiable Composite Rewards
von: Tarek, Mirza Farhan Bin, et al.
Veröffentlicht: (2025)
von: Tarek, Mirza Farhan Bin, et al.
Veröffentlicht: (2025)
Synthetic Blips: Generalizing Synthetic Controls for Dynamic Treatment Effects
von: Agarwal, Anish, et al.
Veröffentlicht: (2022)
von: Agarwal, Anish, et al.
Veröffentlicht: (2022)
Hierarchical Apprenticeship Learning from Imperfect Demonstrations with Evolving Rewards
von: Islam, Md Mirajul, et al.
Veröffentlicht: (2026)
von: Islam, Md Mirajul, et al.
Veröffentlicht: (2026)
Confidence-aware Reward Optimization for Fine-tuning Text-to-Image Models
von: Kim, Kyuyoung, et al.
Veröffentlicht: (2024)
von: Kim, Kyuyoung, et al.
Veröffentlicht: (2024)
Mutual-Taught for Co-adapting Policy and Reward Models
von: Shi, Tianyuan, et al.
Veröffentlicht: (2025)
von: Shi, Tianyuan, et al.
Veröffentlicht: (2025)
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TabImpute: Universal Zero-Shot Imputation for Tabular Data
von: Feitelberg, Jacob, et al.
Veröffentlicht: (2025) -
Distributional Matrix Completion via Nearest Neighbors in the Wasserstein Space
von: Feitelberg, Jacob, et al.
Veröffentlicht: (2024) -
Learning Counterfactual Distributions via Kernel Nearest Neighbors
von: Choi, Kyuseong, et al.
Veröffentlicht: (2024) -
Supervised Kernel Thinning
von: Gong, Albert, et al.
Veröffentlicht: (2024) -
Doubly Robust Inference in Causal Latent Factor Models
von: Abadie, Alberto, et al.
Veröffentlicht: (2024)