Investigating the Interplay of Prioritized Replay and Generalization
Fuente:
arXiv
Salvato in:
| Autori principali: | Panahi, Parham Mohammad, Patterson, Andrew, White, Martha, White, Adam |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A New View on Planning in Online Reinforcement Learning
di: Roice, Kevin, et al.
Pubblicazione: (2024)
di: Roice, Kevin, et al.
Pubblicazione: (2024)
A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
di: Patterson, Andrew, et al.
Pubblicazione: (2021)
di: Patterson, Andrew, et al.
Pubblicazione: (2021)
Forager: a lightweight testbed for continual learning with partial observability in RL
di: Tang, Steven, et al.
Pubblicazione: (2026)
di: Tang, Steven, et al.
Pubblicazione: (2026)
Empirical Design in Reinforcement Learning
di: Patterson, Andrew, et al.
Pubblicazione: (2023)
di: Patterson, Andrew, et al.
Pubblicazione: (2023)
Goal-Space Planning with Subgoal Models
di: Lo, Chunlok, et al.
Pubblicazione: (2022)
di: Lo, Chunlok, et al.
Pubblicazione: (2022)
Deep Reinforcement Learning with Gradient Eligibility Traces
di: Elelimy, Esraa, et al.
Pubblicazione: (2025)
di: Elelimy, Esraa, et al.
Pubblicazione: (2025)
When is Offline Policy Selection Sample Efficient for Reinforcement Learning?
di: Liu, Vincent, et al.
Pubblicazione: (2023)
di: Liu, Vincent, et al.
Pubblicazione: (2023)
Fine-Tuning without Performance Degradation
di: Wang, Han, et al.
Pubblicazione: (2025)
di: Wang, Han, et al.
Pubblicazione: (2025)
Real-Time Recurrent Learning using Trace Units in Reinforcement Learning
di: Elelimy, Esraa, et al.
Pubblicazione: (2024)
di: Elelimy, Esraa, et al.
Pubblicazione: (2024)
Position: Lifetime tuning is incompatible with continual reinforcement learning
di: Mesbahi, Golnaz, et al.
Pubblicazione: (2024)
di: Mesbahi, Golnaz, et al.
Pubblicazione: (2024)
Prioritized Trajectory Replay: A Replay Memory for Data-driven Reinforcement Learning
di: Liu, Jinyi, et al.
Pubblicazione: (2023)
di: Liu, Jinyi, et al.
Pubblicazione: (2023)
Revisiting Mixture Policies in Entropy-Regularized Actor-Critic
di: He, Jiamin, et al.
Pubblicazione: (2026)
di: He, Jiamin, et al.
Pubblicazione: (2026)
Prioritized Replay for RL Post-training
di: Fatemi, Mehdi
Pubblicazione: (2026)
di: Fatemi, Mehdi
Pubblicazione: (2026)
Enhancing LLM Agents for Code Generation with Possibility and Pass-rate Prioritized Experience Replay
di: Chen, Yuyang, et al.
Pubblicazione: (2024)
di: Chen, Yuyang, et al.
Pubblicazione: (2024)
Investigating the Histogram Loss in Regression
di: Imani, Ehsan, et al.
Pubblicazione: (2024)
di: Imani, Ehsan, et al.
Pubblicazione: (2024)
CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
di: Butt, Natasha, et al.
Pubblicazione: (2024)
di: Butt, Natasha, et al.
Pubblicazione: (2024)
DQN Performance with Epsilon Greedy Policies and Prioritized Experience Replay
di: Perkins, Daniel, et al.
Pubblicazione: (2025)
di: Perkins, Daniel, et al.
Pubblicazione: (2025)
Generalized Munchausen Reinforcement Learning using Tsallis KL Divergence
di: Zhu, Lingwei, et al.
Pubblicazione: (2023)
di: Zhu, Lingwei, et al.
Pubblicazione: (2023)
Value Bonuses using Ensemble Errors for Exploration in Reinforcement Learning
di: Wahab, Abdul, et al.
Pubblicazione: (2026)
di: Wahab, Abdul, et al.
Pubblicazione: (2026)
Gradient Iterated Temporal-Difference Learning
di: Vincent, Théo, et al.
Pubblicazione: (2026)
di: Vincent, Théo, et al.
Pubblicazione: (2026)
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers
di: Vasan, Gautham, et al.
Pubblicazione: (2024)
di: Vasan, Gautham, et al.
Pubblicazione: (2024)
Demystifying the Recency Heuristic in Temporal-Difference Learning
di: Daley, Brett, et al.
Pubblicazione: (2024)
di: Daley, Brett, et al.
Pubblicazione: (2024)
Deep Double Q-learning
di: Nagarajan, Prabhat, et al.
Pubblicazione: (2025)
di: Nagarajan, Prabhat, et al.
Pubblicazione: (2025)
Distributions as Actions: A Unified Framework for Diverse Action Spaces
di: He, Jiamin, et al.
Pubblicazione: (2025)
di: He, Jiamin, et al.
Pubblicazione: (2025)
A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning
di: Adkins, Jacob, et al.
Pubblicazione: (2024)
di: Adkins, Jacob, et al.
Pubblicazione: (2024)
D-SPEAR: Dual-Stream Prioritized Experience Adaptive Replay for Stable Reinforcement Learning in Robotic Manipulation
di: Zhang, Yu, et al.
Pubblicazione: (2026)
di: Zhang, Yu, et al.
Pubblicazione: (2026)
Rethinking the Foundations for Continual Reinforcement Learning
di: Elelimy, Esraa, et al.
Pubblicazione: (2025)
di: Elelimy, Esraa, et al.
Pubblicazione: (2025)
The Cross-environment Hyperparameter Setting Benchmark for Reinforcement Learning
di: Patterson, Andrew, et al.
Pubblicazione: (2024)
di: Patterson, Andrew, et al.
Pubblicazione: (2024)
Harnessing Discrete Representations For Continual Reinforcement Learning
di: Meyer, Edan, et al.
Pubblicazione: (2023)
di: Meyer, Edan, et al.
Pubblicazione: (2023)
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values
di: Daley, Brett, et al.
Pubblicazione: (2025)
di: Daley, Brett, et al.
Pubblicazione: (2025)
Better Generative Replay for Continual Federated Learning
di: Qi, Daiqing, et al.
Pubblicazione: (2023)
di: Qi, Daiqing, et al.
Pubblicazione: (2023)
Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers
di: Barron, Joshua, et al.
Pubblicazione: (2025)
di: Barron, Joshua, et al.
Pubblicazione: (2025)
Symmetric Behavior Regularized Policy Optimization
di: Zhu, Lingwei, et al.
Pubblicazione: (2025)
di: Zhu, Lingwei, et al.
Pubblicazione: (2025)
AGaLiTe: Approximate Gated Linear Transformers for Online Reinforcement Learning
di: Pramanik, Subhojeet, et al.
Pubblicazione: (2023)
di: Pramanik, Subhojeet, et al.
Pubblicazione: (2023)
Prioritized Generative Replay
di: Wang, Renhao, et al.
Pubblicazione: (2024)
di: Wang, Renhao, et al.
Pubblicazione: (2024)
Generalized Back-Stepping Experience Replay in Sparse-Reward Environments
di: Lyu, Guwen, et al.
Pubblicazione: (2024)
di: Lyu, Guwen, et al.
Pubblicazione: (2024)
Mitigating Value Hallucination in Dyna Planning via Multistep Predecessor Models
di: Aminmansour, Farzane, et al.
Pubblicazione: (2020)
di: Aminmansour, Farzane, et al.
Pubblicazione: (2020)
Data-Free Generative Replay for Class-Incremental Learning on Imbalanced Data
di: Younis, Sohaib, et al.
Pubblicazione: (2024)
di: Younis, Sohaib, et al.
Pubblicazione: (2024)
VLM-Guided Experience Replay
di: Sharony, Elad, et al.
Pubblicazione: (2026)
di: Sharony, Elad, et al.
Pubblicazione: (2026)
Experience Replay with Random Reshuffling
di: Fujita, Yasuhiro
Pubblicazione: (2025)
di: Fujita, Yasuhiro
Pubblicazione: (2025)
Documenti analoghi
-
A New View on Planning in Online Reinforcement Learning
di: Roice, Kevin, et al.
Pubblicazione: (2024) -
A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
di: Patterson, Andrew, et al.
Pubblicazione: (2021) -
Forager: a lightweight testbed for continual learning with partial observability in RL
di: Tang, Steven, et al.
Pubblicazione: (2026) -
Empirical Design in Reinforcement Learning
di: Patterson, Andrew, et al.
Pubblicazione: (2023) -
Goal-Space Planning with Subgoal Models
di: Lo, Chunlok, et al.
Pubblicazione: (2022)