Are Expressive Models Truly Necessary for Offline RL?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Guan, Niu, Haoyi, Li, Jianxiong, Jiang, Li, Hu, Jianming, Zhan, Xianyuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement Learning
von: Niu, Haoyi, et al.
Veröffentlicht: (2022)
von: Niu, Haoyi, et al.
Veröffentlicht: (2022)
A Comprehensive Survey of Cross-Domain Policy Transfer for Embodied Agents
von: Niu, Haoyi, et al.
Veröffentlicht: (2024)
von: Niu, Haoyi, et al.
Veröffentlicht: (2024)
H2O+: An Improved Framework for Hybrid Offline-and-Online RL with Dynamics Gaps
von: Niu, Haoyi, et al.
Veröffentlicht: (2023)
von: Niu, Haoyi, et al.
Veröffentlicht: (2023)
Skill Expansion and Composition in Parameter Space
von: Liu, Tenglong, et al.
Veröffentlicht: (2025)
von: Liu, Tenglong, et al.
Veröffentlicht: (2025)
Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model
von: Zheng, Yinan, et al.
Veröffentlicht: (2024)
von: Zheng, Yinan, et al.
Veröffentlicht: (2024)
Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
von: Luo, Yu, et al.
Veröffentlicht: (2024)
von: Luo, Yu, et al.
Veröffentlicht: (2024)
Continual Driving Policy Optimization with Closed-Loop Individualized Curricula
von: Niu, Haoyi, et al.
Veröffentlicht: (2023)
von: Niu, Haoyi, et al.
Veröffentlicht: (2023)
Scaling Offline RL via Efficient and Expressive Shortcut Models
von: Espinosa-Dice, Nicolas, et al.
Veröffentlicht: (2025)
von: Espinosa-Dice, Nicolas, et al.
Veröffentlicht: (2025)
Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement Learning
von: Mao, Liyuan, et al.
Veröffentlicht: (2024)
von: Mao, Liyuan, et al.
Veröffentlicht: (2024)
Are Expressive Encoders Necessary for Discrete Graph Generation?
von: Revolinsky, Jay, et al.
Veröffentlicht: (2026)
von: Revolinsky, Jay, et al.
Veröffentlicht: (2026)
xTED: Cross-Domain Adaptation via Diffusion-Based Trajectory Editing
von: Niu, Haoyi, et al.
Veröffentlicht: (2024)
von: Niu, Haoyi, et al.
Veröffentlicht: (2024)
OMPO: A Unified Framework for RL under Policy and Dynamics Shifts
von: Luo, Yu, et al.
Veröffentlicht: (2024)
von: Luo, Yu, et al.
Veröffentlicht: (2024)
Dual Alignment Maximin Optimization for Offline Model-based RL
von: Zhou, Chi, et al.
Veröffentlicht: (2025)
von: Zhou, Chi, et al.
Veröffentlicht: (2025)
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
von: Chen, Zihan, et al.
Veröffentlicht: (2025)
von: Chen, Zihan, et al.
Veröffentlicht: (2025)
Don't "Overthink" Passage Reranking: Is Reasoning Truly Necessary?
von: Jedidi, Nour, et al.
Veröffentlicht: (2025)
von: Jedidi, Nour, et al.
Veröffentlicht: (2025)
DataLight: Offline Data-Driven Traffic Signal Control
von: Zhang, Liang, et al.
Veröffentlicht: (2023)
von: Zhang, Liang, et al.
Veröffentlicht: (2023)
Augmenting Offline RL with Unlabeled Data
von: Wang, Zhao, et al.
Veröffentlicht: (2024)
von: Wang, Zhao, et al.
Veröffentlicht: (2024)
Budgeting Counterfactual for Offline RL
von: Liu, Yao, et al.
Veröffentlicht: (2023)
von: Liu, Yao, et al.
Veröffentlicht: (2023)
Efficient Robotic Policy Learning via Latent Space Backward Planning
von: Liu, Dongxiu, et al.
Veröffentlicht: (2025)
von: Liu, Dongxiu, et al.
Veröffentlicht: (2025)
Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining
von: Cheng, Jie, et al.
Veröffentlicht: (2024)
von: Cheng, Jie, et al.
Veröffentlicht: (2024)
Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs
von: Feng, Zhangying, et al.
Veröffentlicht: (2025)
von: Feng, Zhangying, et al.
Veröffentlicht: (2025)
Scalable Offline Model-Based RL with Action Chunks
von: Park, Kwanyoung, et al.
Veröffentlicht: (2025)
von: Park, Kwanyoung, et al.
Veröffentlicht: (2025)
GAS: Enhancing Reward-Cost Balance of Generative Model-assisted Offline Safe RL
von: Liu, Zifan, et al.
Veröffentlicht: (2026)
von: Liu, Zifan, et al.
Veröffentlicht: (2026)
Robust Policy Expansion for Offline-to-Online RL under Diverse Data Corruption
von: He, Longxiang, et al.
Veröffentlicht: (2025)
von: He, Longxiang, et al.
Veröffentlicht: (2025)
Selective Uncertainty Propagation in Offline RL
von: Krishnamurthy, Sanath Kumar, et al.
Veröffentlicht: (2023)
von: Krishnamurthy, Sanath Kumar, et al.
Veröffentlicht: (2023)
Decoupled Prioritized Resampling for Offline RL
von: Yue, Yang, et al.
Veröffentlicht: (2023)
von: Yue, Yang, et al.
Veröffentlicht: (2023)
DecisionNCE: Embodied Multimodal Representations via Implicit Preference Learning
von: Li, Jianxiong, et al.
Veröffentlicht: (2024)
von: Li, Jianxiong, et al.
Veröffentlicht: (2024)
An Empirical Study on the Effectiveness of Incorporating Offline RL As Online RL Subroutines
von: Su, Jianhai, et al.
Veröffentlicht: (2025)
von: Su, Jianhai, et al.
Veröffentlicht: (2025)
Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?
von: Sun, Guangzhi, et al.
Veröffentlicht: (2025)
von: Sun, Guangzhi, et al.
Veröffentlicht: (2025)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
von: Mark, Max Sobol, et al.
Veröffentlicht: (2024)
von: Mark, Max Sobol, et al.
Veröffentlicht: (2024)
OGBench: Benchmarking Offline Goal-Conditioned RL
von: Park, Seohong, et al.
Veröffentlicht: (2024)
von: Park, Seohong, et al.
Veröffentlicht: (2024)
A Tractable Inference Perspective of Offline RL
von: Liu, Xuejie, et al.
Veröffentlicht: (2023)
von: Liu, Xuejie, et al.
Veröffentlicht: (2023)
Design Considerations in Offline Preference-based RL
von: Agarwal, Alekh, et al.
Veröffentlicht: (2025)
von: Agarwal, Alekh, et al.
Veröffentlicht: (2025)
ODICE: Revealing the Mystery of Distribution Correction Estimation via Orthogonal-gradient Update
von: Mao, Liyuan, et al.
Veröffentlicht: (2024)
von: Mao, Liyuan, et al.
Veröffentlicht: (2024)
Bidirectional-Reachable Hierarchical Reinforcement Learning with Mutually Responsive Policies
von: Luo, Yu, et al.
Veröffentlicht: (2024)
von: Luo, Yu, et al.
Veröffentlicht: (2024)
Are Human-generated Demonstrations Necessary for In-context Learning?
von: Li, Rui, et al.
Veröffentlicht: (2023)
von: Li, Rui, et al.
Veröffentlicht: (2023)
Integrating Domain Knowledge for handling Limited Data in Offline RL
von: Gangopadhyay, Briti, et al.
Veröffentlicht: (2024)
von: Gangopadhyay, Briti, et al.
Veröffentlicht: (2024)
Failure-Aware RL: Reliable Offline-to-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation
von: Li, Huanyu, et al.
Veröffentlicht: (2026)
von: Li, Huanyu, et al.
Veröffentlicht: (2026)
Offline Multi-task Transfer RL with Representational Penalization
von: Bose, Avinandan, et al.
Veröffentlicht: (2024)
von: Bose, Avinandan, et al.
Veröffentlicht: (2024)
Is Value Learning Really the Main Bottleneck in Offline RL?
von: Park, Seohong, et al.
Veröffentlicht: (2024)
von: Park, Seohong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement Learning
von: Niu, Haoyi, et al.
Veröffentlicht: (2022) -
A Comprehensive Survey of Cross-Domain Policy Transfer for Embodied Agents
von: Niu, Haoyi, et al.
Veröffentlicht: (2024) -
H2O+: An Improved Framework for Hybrid Offline-and-Online RL with Dynamics Gaps
von: Niu, Haoyi, et al.
Veröffentlicht: (2023) -
Skill Expansion and Composition in Parameter Space
von: Liu, Tenglong, et al.
Veröffentlicht: (2025) -
Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model
von: Zheng, Yinan, et al.
Veröffentlicht: (2024)