Dual Alignment Maximin Optimization for Offline Model-based RL
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhou, Chi, Luo, Wang, Li, Haoran, Han, Congying, Guo, Tiande, Zhang, Zicheng |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Mitigating Distribution Shift in Model-based Offline RL via Shifts-aware Reward Learning
par: Luo, Wang, et autres
Publié: (2024)
par: Luo, Wang, et autres
Publié: (2024)
Purity Law for Generalizable Neural TSP Solvers
par: Liu, Wenzhao, et autres
Publié: (2025)
par: Liu, Wenzhao, et autres
Publié: (2025)
A Fast Anti-Jamming Cognitive Radar Deployment Algorithm Based on Reinforcement Learning
par: Cai, Wencheng, et autres
Publié: (2025)
par: Cai, Wencheng, et autres
Publié: (2025)
Towards Optimal Adversarial Robust Reinforcement Learning with Infinity Measurement Error
par: Li, Haoran, et autres
Publié: (2025)
par: Li, Haoran, et autres
Publié: (2025)
Towards Optimal Adversarial Robust Q-learning with Bellman Infinity-error
par: Li, Haoran, et autres
Publié: (2024)
par: Li, Haoran, et autres
Publié: (2024)
An Empirical Study on the Effectiveness of Incorporating Offline RL As Online RL Subroutines
par: Su, Jianhai, et autres
Publié: (2025)
par: Su, Jianhai, et autres
Publié: (2025)
Are Expressive Models Truly Necessary for Offline RL?
par: Wang, Guan, et autres
Publié: (2024)
par: Wang, Guan, et autres
Publié: (2024)
On the Tension Between Optimality and Adversarial Robustness in Policy Optimization
par: Li, Haoran, et autres
Publié: (2025)
par: Li, Haoran, et autres
Publié: (2025)
Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining
par: Cheng, Jie, et autres
Publié: (2024)
par: Cheng, Jie, et autres
Publié: (2024)
Design Considerations in Offline Preference-based RL
par: Agarwal, Alekh, et autres
Publié: (2025)
par: Agarwal, Alekh, et autres
Publié: (2025)
Diffusion Models as Optimizers for Efficient Planning in Offline RL
par: Huang, Renming, et autres
Publié: (2024)
par: Huang, Renming, et autres
Publié: (2024)
CROP: Conservative Reward for Model-based Offline Policy Optimization
par: Li, Hao, et autres
Publié: (2023)
par: Li, Hao, et autres
Publié: (2023)
Augmenting Offline RL with Unlabeled Data
par: Wang, Zhao, et autres
Publié: (2024)
par: Wang, Zhao, et autres
Publié: (2024)
Generalized Preference Optimization: A Unified Approach to Offline Alignment
par: Tang, Yunhao, et autres
Publié: (2024)
par: Tang, Yunhao, et autres
Publié: (2024)
Budgeting Counterfactual for Offline RL
par: Liu, Yao, et autres
Publié: (2023)
par: Liu, Yao, et autres
Publié: (2023)
OMG-RL:Offline Model-based Guided Reward Learning for Heparin Treatment
par: Lim, Yooseok, et autres
Publié: (2024)
par: Lim, Yooseok, et autres
Publié: (2024)
Scalable Offline Model-Based RL with Action Chunks
par: Park, Kwanyoung, et autres
Publié: (2025)
par: Park, Kwanyoung, et autres
Publié: (2025)
GAS: Enhancing Reward-Cost Balance of Generative Model-assisted Offline Safe RL
par: Liu, Zifan, et autres
Publié: (2026)
par: Liu, Zifan, et autres
Publié: (2026)
Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies
par: Chen, Jiaqi, et autres
Publié: (2025)
par: Chen, Jiaqi, et autres
Publié: (2025)
Federated Offline Policy Optimization with Dual Regularization
par: Yue, Sheng, et autres
Publié: (2024)
par: Yue, Sheng, et autres
Publié: (2024)
Selective Uncertainty Propagation in Offline RL
par: Krishnamurthy, Sanath Kumar, et autres
Publié: (2023)
par: Krishnamurthy, Sanath Kumar, et autres
Publié: (2023)
Decoupled Prioritized Resampling for Offline RL
par: Yue, Yang, et autres
Publié: (2023)
par: Yue, Yang, et autres
Publié: (2023)
Continual Offline Reinforcement Learning via Diffusion-based Dual Generative Replay
par: Liu, Jinmei, et autres
Publié: (2024)
par: Liu, Jinmei, et autres
Publié: (2024)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
par: Mark, Max Sobol, et autres
Publié: (2024)
par: Mark, Max Sobol, et autres
Publié: (2024)
A Tractable Inference Perspective of Offline RL
par: Liu, Xuejie, et autres
Publié: (2023)
par: Liu, Xuejie, et autres
Publié: (2023)
OGBench: Benchmarking Offline Goal-Conditioned RL
par: Park, Seohong, et autres
Publié: (2024)
par: Park, Seohong, et autres
Publié: (2024)
Scaling Offline RL via Efficient and Expressive Shortcut Models
par: Espinosa-Dice, Nicolas, et autres
Publié: (2025)
par: Espinosa-Dice, Nicolas, et autres
Publié: (2025)
Query-Dependent Prompt Evaluation and Optimization with Offline Inverse RL
par: Sun, Hao, et autres
Publié: (2023)
par: Sun, Hao, et autres
Publié: (2023)
Integrating Domain Knowledge for handling Limited Data in Offline RL
par: Gangopadhyay, Briti, et autres
Publié: (2024)
par: Gangopadhyay, Briti, et autres
Publié: (2024)
Robust Policy Expansion for Offline-to-Online RL under Diverse Data Corruption
par: He, Longxiang, et autres
Publié: (2025)
par: He, Longxiang, et autres
Publié: (2025)
Offline Regularised Reinforcement Learning for Large Language Models Alignment
par: Richemond, Pierre Harvey, et autres
Publié: (2024)
par: Richemond, Pierre Harvey, et autres
Publié: (2024)
Preference-based opponent shaping in differentiable games
par: Qiao, Xinyu, et autres
Publié: (2024)
par: Qiao, Xinyu, et autres
Publié: (2024)
Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
par: Luo, Yu, et autres
Publié: (2024)
par: Luo, Yu, et autres
Publié: (2024)
Offline Imitation Learning with Model-based Reverse Augmentation
par: Shao, Jie-Jing, et autres
Publié: (2024)
par: Shao, Jie-Jing, et autres
Publié: (2024)
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
par: Zhao, Anhao, et autres
Publié: (2026)
par: Zhao, Anhao, et autres
Publié: (2026)
Yes, Q-learning Helps Offline In-Context RL
par: Tarasov, Denis, et autres
Publié: (2025)
par: Tarasov, Denis, et autres
Publié: (2025)
Offline Multi-task Transfer RL with Representational Penalization
par: Bose, Avinandan, et autres
Publié: (2024)
par: Bose, Avinandan, et autres
Publié: (2024)
Is Value Learning Really the Main Bottleneck in Offline RL?
par: Park, Seohong, et autres
Publié: (2024)
par: Park, Seohong, et autres
Publié: (2024)
The Role of Deep Learning Regularizations on Actors in Offline RL
par: Tarasov, Denis, et autres
Publié: (2024)
par: Tarasov, Denis, et autres
Publié: (2024)
Guided Trajectory Generation with Diffusion Models for Offline Model-based Optimization
par: Yun, Taeyoung, et autres
Publié: (2024)
par: Yun, Taeyoung, et autres
Publié: (2024)
Documents similaires
-
Mitigating Distribution Shift in Model-based Offline RL via Shifts-aware Reward Learning
par: Luo, Wang, et autres
Publié: (2024) -
Purity Law for Generalizable Neural TSP Solvers
par: Liu, Wenzhao, et autres
Publié: (2025) -
A Fast Anti-Jamming Cognitive Radar Deployment Algorithm Based on Reinforcement Learning
par: Cai, Wencheng, et autres
Publié: (2025) -
Towards Optimal Adversarial Robust Reinforcement Learning with Infinity Measurement Error
par: Li, Haoran, et autres
Publié: (2025) -
Towards Optimal Adversarial Robust Q-learning with Bellman Infinity-error
par: Li, Haoran, et autres
Publié: (2024)