COPlanner: Plan to Roll Out Conservatively but to Explore Optimistically for Model-Based RL
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xiyao, Zheng, Ruijie, Sun, Yanchao, Jia, Ruonan, Wongkamjan, Wichayaporn, Xu, Huazhe, Huang, Furong |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TACO: Temporal Latent Action-Driven Contrastive Loss for Visual Reinforcement Learning
by: Zheng, Ruijie, et al.
Published: (2023)
by: Zheng, Ruijie, et al.
Published: (2023)
What if Red Can Talk? Dynamic Dialogue Generation Using Large Language Models
by: Nananukul, Navapat, et al.
Published: (2024)
by: Nananukul, Navapat, et al.
Published: (2024)
Adapting Static Fairness to Sequential Decision-Making: Bias Mitigation Strategies towards Equal Long-term Benefit Rate
by: Xu, Yuancheng, et al.
Published: (2023)
by: Xu, Yuancheng, et al.
Published: (2023)
Rethinking Adversarial Policies: A Generalized Attack Formulation and Provable Defense in RL
by: Liu, Xiangyu, et al.
Published: (2023)
by: Liu, Xiangyu, et al.
Published: (2023)
World Models with Hints of Large Language Models for Goal Achieving
by: Liu, Zeyuan, et al.
Published: (2024)
by: Liu, Zeyuan, et al.
Published: (2024)
Beyond Worst-case Attacks: Robust RL with Adaptive Defense via Non-dominated Policies
by: Liu, Xiangyu, et al.
Published: (2024)
by: Liu, Xiangyu, et al.
Published: (2024)
SOMBRL: Scalable and Optimistic Model-Based RL
by: Sukhija, Bhavya, et al.
Published: (2025)
by: Sukhija, Bhavya, et al.
Published: (2025)
Premier-TACO is a Few-Shot Policy Learner: Pretraining Multitask Representation via Temporal Action-Driven Contrastive Loss
by: Zheng, Ruijie, et al.
Published: (2024)
by: Zheng, Ruijie, et al.
Published: (2024)
ACE : Off-Policy Actor-Critic with Causality-Aware Entropy Regularization
by: Ji, Tianying, et al.
Published: (2024)
by: Ji, Tianying, et al.
Published: (2024)
Game-Theoretic Robust Reinforcement Learning Handles Temporally-Coupled Perturbations
by: Liang, Yongyuan, et al.
Published: (2023)
by: Liang, Yongyuan, et al.
Published: (2023)
Should I Trust You? Detecting Deception in Negotiations using Counterfactual RL
by: Wongkamjan, Wichayaporn, et al.
Published: (2025)
by: Wongkamjan, Wichayaporn, et al.
Published: (2025)
DrM: Mastering Visual Reinforcement Learning through Dormant Ratio Minimization
by: Xu, Guowei, et al.
Published: (2023)
by: Xu, Guowei, et al.
Published: (2023)
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL
by: Luo, Qin-Wen, et al.
Published: (2024)
by: Luo, Qin-Wen, et al.
Published: (2024)
A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control
by: Kang, Zilin, et al.
Published: (2025)
by: Kang, Zilin, et al.
Published: (2025)
OMPO: A Unified Framework for RL under Policy and Dynamics Shifts
by: Luo, Yu, et al.
Published: (2024)
by: Luo, Yu, et al.
Published: (2024)
PRISE: LLM-Style Sequence Compression for Learning Temporal Action Abstractions in Control
by: Zheng, Ruijie, et al.
Published: (2024)
by: Zheng, Ruijie, et al.
Published: (2024)
Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
by: Luo, Yu, et al.
Published: (2024)
by: Luo, Yu, et al.
Published: (2024)
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
by: Xu, Yuancheng, et al.
Published: (2024)
by: Xu, Yuancheng, et al.
Published: (2024)
Personalized Help for Optimizing Low-Skilled Users' Strategy
by: Gu, Feng, et al.
Published: (2024)
by: Gu, Feng, et al.
Published: (2024)
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
by: Wang, Xiyao, et al.
Published: (2025)
by: Wang, Xiyao, et al.
Published: (2025)
Optimistic ε-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning
by: Zhang, Ruoning, et al.
Published: (2025)
by: Zhang, Ruoning, et al.
Published: (2025)
Agentic Critical Training
by: Liu, Weize, et al.
Published: (2026)
by: Liu, Weize, et al.
Published: (2026)
On the Evaluation of Generative Robotic Simulations
by: Chen, Feng, et al.
Published: (2024)
by: Chen, Feng, et al.
Published: (2024)
Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
by: Zhai, Yuanzhao, et al.
Published: (2024)
by: Zhai, Yuanzhao, et al.
Published: (2024)
RollArt: Scaling Agentic RL Training via Disaggregated Infrastructure
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy Constraints
by: Kang, Zilin, et al.
Published: (2025)
by: Kang, Zilin, et al.
Published: (2025)
Efficient Model-Based Reinforcement Learning Through Optimistic Thompson Sampling
by: Bayrooti, Jasmine, et al.
Published: (2024)
by: Bayrooti, Jasmine, et al.
Published: (2024)
RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
Optimistic Task Inference for Behavior Foundation Models
by: Rupf, Thomas, et al.
Published: (2025)
by: Rupf, Thomas, et al.
Published: (2025)
Optimistically Optimistic Exploration for Provably Efficient Infinite-Horizon Reinforcement and Imitation Learning
by: Moulin, Antoine, et al.
Published: (2025)
by: Moulin, Antoine, et al.
Published: (2025)
DittoGym: Learning to Control Soft Shape-Shifting Robots
by: Huang, Suning, et al.
Published: (2024)
by: Huang, Suning, et al.
Published: (2024)
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
by: Wang, Xiyao, et al.
Published: (2024)
by: Wang, Xiyao, et al.
Published: (2024)
TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning
by: Wu, Jinyang, et al.
Published: (2025)
by: Wu, Jinyang, et al.
Published: (2025)
Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
by: Wang, Xiyao, et al.
Published: (2024)
by: Wang, Xiyao, et al.
Published: (2024)
Optimistic Policy Regularization
by: Pham, Mai, et al.
Published: (2026)
by: Pham, Mai, et al.
Published: (2026)
Failure-Aware RL: Reliable Offline-to-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation
by: Li, Huanyu, et al.
Published: (2026)
by: Li, Huanyu, et al.
Published: (2026)
Optimistic World Models: Efficient Exploration in Model-Based Deep Reinforcement Learning
by: Mete, Akshay, et al.
Published: (2026)
by: Mete, Akshay, et al.
Published: (2026)
Optimistic Information Directed Sampling
by: Neu, Gergely, et al.
Published: (2024)
by: Neu, Gergely, et al.
Published: (2024)
RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning
by: Lei, Kun, et al.
Published: (2025)
by: Lei, Kun, et al.
Published: (2025)
Conservative DDPG -- Pessimistic RL without Ensemble
by: Soffair, Nitsan, et al.
Published: (2024)
by: Soffair, Nitsan, et al.
Published: (2024)
Similar Items
-
TACO: Temporal Latent Action-Driven Contrastive Loss for Visual Reinforcement Learning
by: Zheng, Ruijie, et al.
Published: (2023) -
What if Red Can Talk? Dynamic Dialogue Generation Using Large Language Models
by: Nananukul, Navapat, et al.
Published: (2024) -
Adapting Static Fairness to Sequential Decision-Making: Bias Mitigation Strategies towards Equal Long-term Benefit Rate
by: Xu, Yuancheng, et al.
Published: (2023) -
Rethinking Adversarial Policies: A Generalized Attack Formulation and Provable Defense in RL
by: Liu, Xiangyu, et al.
Published: (2023) -
World Models with Hints of Large Language Models for Goal Achieving
by: Liu, Zeyuan, et al.
Published: (2024)