Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhai, Yuanzhao, Li, Yiying, Gao, Zijian, Gong, Xudong, Xu, Kele, Feng, Dawei, Bo, Ding, Wang, Huaimin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
von: Zhai, Yuanzhao, et al.
Veröffentlicht: (2023)
von: Zhai, Yuanzhao, et al.
Veröffentlicht: (2023)
Online Self-Preferring Language Models
von: Zhai, Yuanzhao, et al.
Veröffentlicht: (2024)
von: Zhai, Yuanzhao, et al.
Veröffentlicht: (2024)
Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models
von: Zhai, Yuanzhao, et al.
Veröffentlicht: (2024)
von: Zhai, Yuanzhao, et al.
Veröffentlicht: (2024)
Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees
von: Ganguly, Sourav, et al.
Veröffentlicht: (2026)
von: Ganguly, Sourav, et al.
Veröffentlicht: (2026)
MAny: Merge Anything for Multimodal Continual Instruction Tuning
von: Gao, Zijian, et al.
Veröffentlicht: (2026)
von: Gao, Zijian, et al.
Veröffentlicht: (2026)
Scaffold-Conditioned Preference Triplets for Controllable Molecular Optimization with Large Language Models
von: Xiong, Yi, et al.
Veröffentlicht: (2026)
von: Xiong, Yi, et al.
Veröffentlicht: (2026)
Diverse Randomized Value Functions: A Provably Pessimistic Approach for Offline Reinforcement Learning
von: Yu, Xudong, et al.
Veröffentlicht: (2024)
von: Yu, Xudong, et al.
Veröffentlicht: (2024)
Hyperparameter Tuning Through Pessimistic Bilevel Optimization
von: Ustun, Meltem Apaydin, et al.
Veröffentlicht: (2024)
von: Ustun, Meltem Apaydin, et al.
Veröffentlicht: (2024)
Pessimistic Causal Reinforcement Learning with Mediators for Confounded Offline Data
von: Wang, Danyang, et al.
Veröffentlicht: (2024)
von: Wang, Danyang, et al.
Veröffentlicht: (2024)
Pessimistic Off-Policy Optimization for Learning to Rank
von: Cief, Matej, et al.
Veröffentlicht: (2022)
von: Cief, Matej, et al.
Veröffentlicht: (2022)
Pessimistic Backward Policy for GFlowNets
von: Jang, Hyosoon, et al.
Veröffentlicht: (2024)
von: Jang, Hyosoon, et al.
Veröffentlicht: (2024)
Pessimistic Value Iteration for Multi-Task Data Sharing in Offline Reinforcement Learning
von: Bai, Chenjia, et al.
Veröffentlicht: (2024)
von: Bai, Chenjia, et al.
Veröffentlicht: (2024)
Pessimistic Risk-Aware Policy Learning in Contextual Bandits
von: Wan, Yilong, et al.
Veröffentlicht: (2026)
von: Wan, Yilong, et al.
Veröffentlicht: (2026)
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
von: Zhong, Han, et al.
Veröffentlicht: (2021)
von: Zhong, Han, et al.
Veröffentlicht: (2021)
Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning
von: Ding, Zihan, et al.
Veröffentlicht: (2024)
von: Ding, Zihan, et al.
Veröffentlicht: (2024)
Feasibility-Aware Pessimistic Estimation: Toward Long-Horizon Safety in Offline RL
von: Tao, Zhikun
Veröffentlicht: (2025)
von: Tao, Zhikun
Veröffentlicht: (2025)
Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement Learning
von: Di, Qiwei, et al.
Veröffentlicht: (2023)
von: Di, Qiwei, et al.
Veröffentlicht: (2023)
Affective Music Recommendation: A Rollout-Based World Model for Offline Preference Optimization
von: Chan, Audrey, et al.
Veröffentlicht: (2026)
von: Chan, Audrey, et al.
Veröffentlicht: (2026)
COPR: Continual Learning Human Preference through Optimal Policy Regularization
von: Zhang, Han, et al.
Veröffentlicht: (2023)
von: Zhang, Han, et al.
Veröffentlicht: (2023)
Optimistic Policy Regularization
von: Pham, Mai, et al.
Veröffentlicht: (2026)
von: Pham, Mai, et al.
Veröffentlicht: (2026)
Learning a Pessimistic Reward Model in RLHF
von: Xu, Yinglun, et al.
Veröffentlicht: (2025)
von: Xu, Yinglun, et al.
Veröffentlicht: (2025)
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR
von: Wang, Tao, et al.
Veröffentlicht: (2026)
von: Wang, Tao, et al.
Veröffentlicht: (2026)
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL
von: Luo, Qin-Wen, et al.
Veröffentlicht: (2024)
von: Luo, Qin-Wen, et al.
Veröffentlicht: (2024)
Selective Rollout: Mid-Trajectory Termination for Multi-Sample Agent RL
von: Zhai, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zhai, Zhiyuan, et al.
Veröffentlicht: (2026)
Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning
von: Sakhi, Otmane, et al.
Veröffentlicht: (2024)
von: Sakhi, Otmane, et al.
Veröffentlicht: (2024)
Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models
von: Cohen, Lior, et al.
Veröffentlicht: (2026)
von: Cohen, Lior, et al.
Veröffentlicht: (2026)
CROP: Conservative Reward for Model-based Offline Policy Optimization
von: Li, Hao, et al.
Veröffentlicht: (2023)
von: Li, Hao, et al.
Veröffentlicht: (2023)
Optimistic Multi-Agent Policy Gradient
von: Zhao, Wenshuai, et al.
Veröffentlicht: (2023)
von: Zhao, Wenshuai, et al.
Veröffentlicht: (2023)
POLAR: A Pessimistic Model-based Policy Learning Algorithm for Dynamic Treatment Regimes
von: Zhang, Ruijia, et al.
Veröffentlicht: (2025)
von: Zhang, Ruijia, et al.
Veröffentlicht: (2025)
Breaking the Curse of Repulsion: Optimistic Distributionally Robust Policy Optimization for Off-Policy Generative Recommendation
von: Jiang, Jie, et al.
Veröffentlicht: (2026)
von: Jiang, Jie, et al.
Veröffentlicht: (2026)
Knowledgeable Agents by Offline Reinforcement Learning from Large Language Model Rollouts
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2024)
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2024)
Fat-to-Thin Policy Optimization: Offline RL with Sparse Policies
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning
von: Mundada, Gagan, et al.
Veröffentlicht: (2026)
von: Mundada, Gagan, et al.
Veröffentlicht: (2026)
Your Offline Policy is Not Trustworthy: Bilevel Reinforcement Learning for Sequential Portfolio Optimization
von: Yuan, Haochen, et al.
Veröffentlicht: (2025)
von: Yuan, Haochen, et al.
Veröffentlicht: (2025)
DyDiff: Long-Horizon Rollout via Dynamics Diffusion for Offline Reinforcement Learning
von: Zhao, Hanye, et al.
Veröffentlicht: (2024)
von: Zhao, Hanye, et al.
Veröffentlicht: (2024)
Off-Policy Safe Reinforcement Learning with Constrained Optimistic Exploration
von: Li, Guopeng, et al.
Veröffentlicht: (2026)
von: Li, Guopeng, et al.
Veröffentlicht: (2026)
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
von: Zhang, Xuechen, et al.
Veröffentlicht: (2025)
von: Zhang, Xuechen, et al.
Veröffentlicht: (2025)
Optimistic Dual Averaging Unifies Modern Optimizers
von: Pethick, Thomas, et al.
Veröffentlicht: (2026)
von: Pethick, Thomas, et al.
Veröffentlicht: (2026)
DiffPoGAN: Diffusion Policies with Generative Adversarial Networks for Offline Reinforcement Learning
von: Hu, Xuemin, et al.
Veröffentlicht: (2024)
von: Hu, Xuemin, et al.
Veröffentlicht: (2024)
Inference Time Policy Optimization for Offline RL with Differentiable World Models
von: Deb, Rohan, et al.
Veröffentlicht: (2026)
von: Deb, Rohan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
von: Zhai, Yuanzhao, et al.
Veröffentlicht: (2023) -
Online Self-Preferring Language Models
von: Zhai, Yuanzhao, et al.
Veröffentlicht: (2024) -
Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models
von: Zhai, Yuanzhao, et al.
Veröffentlicht: (2024) -
Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees
von: Ganguly, Sourav, et al.
Veröffentlicht: (2026) -
MAny: Merge Anything for Multimodal Continual Instruction Tuning
von: Gao, Zijian, et al.
Veröffentlicht: (2026)