A New View on Planning in Online Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Roice, Kevin, Panahi, Parham Mohammad, Jordan, Scott M., White, Adam, White, Martha |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Goal-Space Planning with Subgoal Models
by: Lo, Chunlok, et al.
Published: (2022)
by: Lo, Chunlok, et al.
Published: (2022)
Investigating the Interplay of Prioritized Replay and Generalization
by: Panahi, Parham Mohammad, et al.
Published: (2024)
by: Panahi, Parham Mohammad, et al.
Published: (2024)
A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
by: Patterson, Andrew, et al.
Published: (2021)
by: Patterson, Andrew, et al.
Published: (2021)
Empirical Design in Reinforcement Learning
by: Patterson, Andrew, et al.
Published: (2023)
by: Patterson, Andrew, et al.
Published: (2023)
Real-Time Recurrent Learning using Trace Units in Reinforcement Learning
by: Elelimy, Esraa, et al.
Published: (2024)
by: Elelimy, Esraa, et al.
Published: (2024)
Forager: a lightweight testbed for continual learning with partial observability in RL
by: Tang, Steven, et al.
Published: (2026)
by: Tang, Steven, et al.
Published: (2026)
Deep Reinforcement Learning with Gradient Eligibility Traces
by: Elelimy, Esraa, et al.
Published: (2025)
by: Elelimy, Esraa, et al.
Published: (2025)
Fine-Tuning without Performance Degradation
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
AGaLiTe: Approximate Gated Linear Transformers for Online Reinforcement Learning
by: Pramanik, Subhojeet, et al.
Published: (2023)
by: Pramanik, Subhojeet, et al.
Published: (2023)
A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning
by: Adkins, Jacob, et al.
Published: (2024)
by: Adkins, Jacob, et al.
Published: (2024)
Value Bonuses using Ensemble Errors for Exploration in Reinforcement Learning
by: Wahab, Abdul, et al.
Published: (2026)
by: Wahab, Abdul, et al.
Published: (2026)
Rethinking the Foundations for Continual Reinforcement Learning
by: Elelimy, Esraa, et al.
Published: (2025)
by: Elelimy, Esraa, et al.
Published: (2025)
Position: Lifetime tuning is incompatible with continual reinforcement learning
by: Mesbahi, Golnaz, et al.
Published: (2024)
by: Mesbahi, Golnaz, et al.
Published: (2024)
Harnessing Discrete Representations For Continual Reinforcement Learning
by: Meyer, Edan, et al.
Published: (2023)
by: Meyer, Edan, et al.
Published: (2023)
Generalized Munchausen Reinforcement Learning using Tsallis KL Divergence
by: Zhu, Lingwei, et al.
Published: (2023)
by: Zhu, Lingwei, et al.
Published: (2023)
When is Offline Policy Selection Sample Efficient for Reinforcement Learning?
by: Liu, Vincent, et al.
Published: (2023)
by: Liu, Vincent, et al.
Published: (2023)
Gradient Iterated Temporal-Difference Learning
by: Vincent, Théo, et al.
Published: (2026)
by: Vincent, Théo, et al.
Published: (2026)
A Systematic Investigation of The RL-Jailbreaker in LLMs
by: Mohammedalamen, Montaser, et al.
Published: (2026)
by: Mohammedalamen, Montaser, et al.
Published: (2026)
Revisiting Mixture Policies in Entropy-Regularized Actor-Critic
by: He, Jiamin, et al.
Published: (2026)
by: He, Jiamin, et al.
Published: (2026)
VISTA: A Panoramic View of Neural Representations
by: White, Tom
Published: (2024)
by: White, Tom
Published: (2024)
Demystifying the Recency Heuristic in Temporal-Difference Learning
by: Daley, Brett, et al.
Published: (2024)
by: Daley, Brett, et al.
Published: (2024)
Position: Benchmarking is Limited in Reinforcement Learning Research
by: Jordan, Scott M., et al.
Published: (2024)
by: Jordan, Scott M., et al.
Published: (2024)
Mitigating Value Hallucination in Dyna Planning via Multistep Predecessor Models
by: Aminmansour, Farzane, et al.
Published: (2020)
by: Aminmansour, Farzane, et al.
Published: (2020)
Distributions as Actions: A Unified Framework for Diverse Action Spaces
by: He, Jiamin, et al.
Published: (2025)
by: He, Jiamin, et al.
Published: (2025)
Continual Reinforcement Learning by Planning with Online World Models
by: Liu, Zichen, et al.
Published: (2025)
by: Liu, Zichen, et al.
Published: (2025)
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values
by: Daley, Brett, et al.
Published: (2025)
by: Daley, Brett, et al.
Published: (2025)
Human-Inspired Multi-Level Reinforcement Learning
by: Wu, Mingkang, et al.
Published: (2025)
by: Wu, Mingkang, et al.
Published: (2025)
Deep Double Q-learning
by: Nagarajan, Prabhat, et al.
Published: (2025)
by: Nagarajan, Prabhat, et al.
Published: (2025)
Performance Optimization of Ratings-Based Reinforcement Learning
by: Rose, Evelyn, et al.
Published: (2025)
by: Rose, Evelyn, et al.
Published: (2025)
Rating-based Reinforcement Learning
by: White, Devin, et al.
Published: (2023)
by: White, Devin, et al.
Published: (2023)
Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models
by: Jajoo, Pranaya, et al.
Published: (2026)
by: Jajoo, Pranaya, et al.
Published: (2026)
Online Reinforcement Learning in Non-Stationary Context-Driven Environments
by: Hamadanian, Pouya, et al.
Published: (2023)
by: Hamadanian, Pouya, et al.
Published: (2023)
HypergraphFormer: Learning Hypergraphs from LLMs for Editable Floor Plan Generation
by: Klimenko, Nikita, et al.
Published: (2026)
by: Klimenko, Nikita, et al.
Published: (2026)
A Study of Plasticity Loss in On-Policy Deep Reinforcement Learning
by: Juliani, Arthur, et al.
Published: (2024)
by: Juliani, Arthur, et al.
Published: (2024)
A New Perspective on Transformers in Online Reinforcement Learning for Continuous Control
by: Kachaev, Nikita, et al.
Published: (2025)
by: Kachaev, Nikita, et al.
Published: (2025)
On Zero-Shot Reinforcement Learning
by: Jeen, Scott
Published: (2025)
by: Jeen, Scott
Published: (2025)
Reinforcement Learning: An Overview
by: Murphy, Kevin
Published: (2024)
by: Murphy, Kevin
Published: (2024)
Symmetric Behavior Regularized Policy Optimization
by: Zhu, Lingwei, et al.
Published: (2025)
by: Zhu, Lingwei, et al.
Published: (2025)
Investigating Action Encodings in Recurrent Neural Networks in Reinforcement Learning
by: Schlegel, Matthew, et al.
Published: (2026)
by: Schlegel, Matthew, et al.
Published: (2026)
Environment-Aware Transfer Reinforcement Learning for Sustainable Beam Selection
by: Salami, Dariush, et al.
Published: (2025)
by: Salami, Dariush, et al.
Published: (2025)
Similar Items
-
Goal-Space Planning with Subgoal Models
by: Lo, Chunlok, et al.
Published: (2022) -
Investigating the Interplay of Prioritized Replay and Generalization
by: Panahi, Parham Mohammad, et al.
Published: (2024) -
A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
by: Patterson, Andrew, et al.
Published: (2021) -
Empirical Design in Reinforcement Learning
by: Patterson, Andrew, et al.
Published: (2023) -
Real-Time Recurrent Learning using Trace Units in Reinforcement Learning
by: Elelimy, Esraa, et al.
Published: (2024)