Efficient Offline Reinforcement Learning: First Imitate, then Improve
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jelley, Adam, McInroe, Trevor, Devlin, Sam, Storkey, Amos |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Planning to Go Out-of-Distribution in Offline-to-Online Reinforcement Learning
von: McInroe, Trevor, et al.
Veröffentlicht: (2023)
von: McInroe, Trevor, et al.
Veröffentlicht: (2023)
Object-Centric World Models from Few-Shot Annotations for Sample-Efficient Reinforcement Learning
von: Zhang, Weipu, et al.
Veröffentlicht: (2025)
von: Zhang, Weipu, et al.
Veröffentlicht: (2025)
PixelBrax: Learning Continuous Control from Pixels End-to-End on the GPU
von: McInroe, Trevor, et al.
Veröffentlicht: (2025)
von: McInroe, Trevor, et al.
Veröffentlicht: (2025)
LLM-Personalize: Aligning LLM Planners with Human Preferences via Reinforced Self-Training for Housekeeping Robots
von: Han, Dongge, et al.
Veröffentlicht: (2024)
von: Han, Dongge, et al.
Veröffentlicht: (2024)
Multi-Horizon Representations with Hierarchical Forward Models for Reinforcement Learning
von: McInroe, Trevor, et al.
Veröffentlicht: (2022)
von: McInroe, Trevor, et al.
Veröffentlicht: (2022)
Aligning Agents like Large Language Models
von: Jelley, Adam, et al.
Veröffentlicht: (2024)
von: Jelley, Adam, et al.
Veröffentlicht: (2024)
Forgetting is Everywhere
von: Sanati, Ben, et al.
Veröffentlicht: (2025)
von: Sanati, Ben, et al.
Veröffentlicht: (2025)
Enhancing Tactile-based Reinforcement Learning for Robotic Control
von: Miller, Elle, et al.
Veröffentlicht: (2025)
von: Miller, Elle, et al.
Veröffentlicht: (2025)
CODA: Coordination via On-Policy Diffusion for Multi-Agent Offline Reinforcement Learning
von: Hedman, Marcel, et al.
Veröffentlicht: (2026)
von: Hedman, Marcel, et al.
Veröffentlicht: (2026)
Terra Nova: A Comprehensive Challenge Environment for Intelligent Agents
von: McInroe, Trevor
Veröffentlicht: (2025)
von: McInroe, Trevor
Veröffentlicht: (2025)
Rationality Measurement and Theory for Reinforcement Learning Agents
von: Qian, Kejiang, et al.
Veröffentlicht: (2026)
von: Qian, Kejiang, et al.
Veröffentlicht: (2026)
Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning
von: Garcin, Samuel, et al.
Veröffentlicht: (2025)
von: Garcin, Samuel, et al.
Veröffentlicht: (2025)
Diffusion for World Modeling: Visual Details Matter in Atari
von: Alonso, Eloi, et al.
Veröffentlicht: (2024)
von: Alonso, Eloi, et al.
Veröffentlicht: (2024)
Chunking: Continual Learning is not just about Distribution Shift
von: Lee, Thomas L., et al.
Veröffentlicht: (2023)
von: Lee, Thomas L., et al.
Veröffentlicht: (2023)
Noisy Early Stopping for Noisy Labels
von: Toner, William, et al.
Veröffentlicht: (2024)
von: Toner, William, et al.
Veröffentlicht: (2024)
Label Noise: Correcting the Forward-Correction
von: Toner, William, et al.
Veröffentlicht: (2023)
von: Toner, William, et al.
Veröffentlicht: (2023)
Assistax: A Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics
von: Hinckeldey, Leonard, et al.
Veröffentlicht: (2025)
von: Hinckeldey, Leonard, et al.
Veröffentlicht: (2025)
roto 2.0: The Robot Tactile Olympiad
von: Miller, Elle, et al.
Veröffentlicht: (2026)
von: Miller, Elle, et al.
Veröffentlicht: (2026)
Guided Data Augmentation for Offline Reinforcement Learning and Imitation Learning
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
Approximate Bayesian Class-Conditional Models under Continuous Representation Shift
von: Lee, Thomas L., et al.
Veröffentlicht: (2023)
von: Lee, Thomas L., et al.
Veröffentlicht: (2023)
Adversarial robustness of VAEs through the lens of local geometry
von: Khan, Asif, et al.
Veröffentlicht: (2022)
von: Khan, Asif, et al.
Veröffentlicht: (2022)
From Novelty to Imitation: Self-Distilled Rewards for Offline Reinforcement Learning
von: Chaudhary, Gaurav, et al.
Veröffentlicht: (2025)
von: Chaudhary, Gaurav, et al.
Veröffentlicht: (2025)
Adapting Time Series Foundation Models through Data Mixtures
von: Lee, Thomas L., et al.
Veröffentlicht: (2026)
von: Lee, Thomas L., et al.
Veröffentlicht: (2026)
Offline Imitation Learning with Variational Counterfactual Reasoning
von: He, Bowei, et al.
Veröffentlicht: (2023)
von: He, Bowei, et al.
Veröffentlicht: (2023)
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
von: Gupta, Akash, et al.
Veröffentlicht: (2025)
von: Gupta, Akash, et al.
Veröffentlicht: (2025)
Kalman Linear Attention: Parallel Bayesian Filtering For Efficient Language Modelling and State Tracking
von: Shaj, Vaisakh, et al.
Veröffentlicht: (2026)
von: Shaj, Vaisakh, et al.
Veröffentlicht: (2026)
Markov Balance Satisfaction Improves Performance in Strictly Batch Offline Imitation Learning
von: Agrawal, Rishabh, et al.
Veröffentlicht: (2024)
von: Agrawal, Rishabh, et al.
Veröffentlicht: (2024)
Few-Shot Learning with Class Imbalance
von: Ochal, Mateusz, et al.
Veröffentlicht: (2021)
von: Ochal, Mateusz, et al.
Veröffentlicht: (2021)
Hyperparameter Selection in Continual Learning
von: Lee, Thomas L., et al.
Veröffentlicht: (2024)
von: Lee, Thomas L., et al.
Veröffentlicht: (2024)
Offline Imitation Learning by Controlling the Effective Planning Horizon
von: Ahn, Hee-Jun, et al.
Veröffentlicht: (2024)
von: Ahn, Hee-Jun, et al.
Veröffentlicht: (2024)
Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning
von: Huang, Kevin, et al.
Veröffentlicht: (2025)
von: Huang, Kevin, et al.
Veröffentlicht: (2025)
Flexible Blood Glucose Control: Offline Reinforcement Learning from Human Feedback
von: Emerson, Harry, et al.
Veröffentlicht: (2025)
von: Emerson, Harry, et al.
Veröffentlicht: (2025)
Visual Encoders for Data-Efficient Imitation Learning in Modern Video Games
von: Schäfer, Lukas, et al.
Veröffentlicht: (2023)
von: Schäfer, Lukas, et al.
Veröffentlicht: (2023)
DITTO: Offline Imitation Learning with World Models
von: DeMoss, Branton, et al.
Veröffentlicht: (2023)
von: DeMoss, Branton, et al.
Veröffentlicht: (2023)
Robust Offline Imitation Learning from Diverse Auxiliary Data
von: Ghosh, Udita, et al.
Veröffentlicht: (2024)
von: Ghosh, Udita, et al.
Veröffentlicht: (2024)
Zero-Shot Offline Imitation Learning via Optimal Transport
von: Rupf, Thomas, et al.
Veröffentlicht: (2024)
von: Rupf, Thomas, et al.
Veröffentlicht: (2024)
Improving Offline Reinforcement Learning with Inaccurate Simulators
von: Hou, Yiwen, et al.
Veröffentlicht: (2024)
von: Hou, Yiwen, et al.
Veröffentlicht: (2024)
Offline Imitation Learning with Model-based Reverse Augmentation
von: Shao, Jie-Jing, et al.
Veröffentlicht: (2024)
von: Shao, Jie-Jing, et al.
Veröffentlicht: (2024)
How to Leverage Diverse Demonstrations in Offline Imitation Learning
von: Yue, Sheng, et al.
Veröffentlicht: (2024)
von: Yue, Sheng, et al.
Veröffentlicht: (2024)
Deterministic Uncertainty Propagation for Improved Model-Based Offline Reinforcement Learning
von: Akgül, Abdullah, et al.
Veröffentlicht: (2024)
von: Akgül, Abdullah, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Planning to Go Out-of-Distribution in Offline-to-Online Reinforcement Learning
von: McInroe, Trevor, et al.
Veröffentlicht: (2023) -
Object-Centric World Models from Few-Shot Annotations for Sample-Efficient Reinforcement Learning
von: Zhang, Weipu, et al.
Veröffentlicht: (2025) -
PixelBrax: Learning Continuous Control from Pixels End-to-End on the GPU
von: McInroe, Trevor, et al.
Veröffentlicht: (2025) -
LLM-Personalize: Aligning LLM Planners with Human Preferences via Reinforced Self-Training for Housekeeping Robots
von: Han, Dongge, et al.
Veröffentlicht: (2024) -
Multi-Horizon Representations with Hierarchical Forward Models for Reinforcement Learning
von: McInroe, Trevor, et al.
Veröffentlicht: (2022)