Deep Double Q-learning
Fuente:
arXiv
Saved in:
| Main Authors: | Nagarajan, Prabhat, White, Martha, Machado, Marlos C. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values
by: Daley, Brett, et al.
Published: (2025)
by: Daley, Brett, et al.
Published: (2025)
Demystifying the Recency Heuristic in Temporal-Difference Learning
by: Daley, Brett, et al.
Published: (2024)
by: Daley, Brett, et al.
Published: (2024)
Deep Reinforcement Learning with Gradient Eligibility Traces
by: Elelimy, Esraa, et al.
Published: (2025)
by: Elelimy, Esraa, et al.
Published: (2025)
When is Offline Policy Selection Sample Efficient for Reinforcement Learning?
by: Liu, Vincent, et al.
Published: (2023)
by: Liu, Vincent, et al.
Published: (2023)
Harnessing Discrete Representations For Continual Reinforcement Learning
by: Meyer, Edan, et al.
Published: (2023)
by: Meyer, Edan, et al.
Published: (2023)
AGaLiTe: Approximate Gated Linear Transformers for Online Reinforcement Learning
by: Pramanik, Subhojeet, et al.
Published: (2023)
by: Pramanik, Subhojeet, et al.
Published: (2023)
The Laplacian Keyboard: Beyond the Linear Span
by: Chandrasekar, Siddarth, et al.
Published: (2026)
by: Chandrasekar, Siddarth, et al.
Published: (2026)
Proper Laplacian Representation Learning
by: Gomez, Diego, et al.
Published: (2023)
by: Gomez, Diego, et al.
Published: (2023)
Double Successive Over-Relaxation Q-Learning with an Extension to Deep Reinforcement Learning
by: R, Shreyas S
Published: (2024)
by: R, Shreyas S
Published: (2024)
Averaging $n$-step Returns Reduces Variance in Reinforcement Learning
by: Daley, Brett, et al.
Published: (2024)
by: Daley, Brett, et al.
Published: (2024)
Learning Tennis Strategy Through Curriculum-Based Dueling Double Deep Q-Networks
by: Mohan, Vishnu
Published: (2025)
by: Mohan, Vishnu
Published: (2025)
The Cell Must Go On: Agar.io for Continual Reinforcement Learning
by: Mohamed, Mohamed A., et al.
Published: (2025)
by: Mohamed, Mohamed A., et al.
Published: (2025)
ARDDQN: Attention Recurrent Double Deep Q-Network for UAV Coverage Path Planning and Data Harvesting
by: Kumar, Praveen, et al.
Published: (2024)
by: Kumar, Praveen, et al.
Published: (2024)
Fine-Tuning without Performance Degradation
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
by: Patterson, Andrew, et al.
Published: (2021)
by: Patterson, Andrew, et al.
Published: (2021)
Deep Reinforcement Learning with Spiking Q-learning
by: Chen, Ding, et al.
Published: (2022)
by: Chen, Ding, et al.
Published: (2022)
Real-Time Recurrent Learning using Trace Units in Reinforcement Learning
by: Elelimy, Esraa, et al.
Published: (2024)
by: Elelimy, Esraa, et al.
Published: (2024)
Empirical Design in Reinforcement Learning
by: Patterson, Andrew, et al.
Published: (2023)
by: Patterson, Andrew, et al.
Published: (2023)
Is Q-learning an Ill-posed Problem?
by: Wissmann, Philipp, et al.
Published: (2025)
by: Wissmann, Philipp, et al.
Published: (2025)
MinMaxMin $Q$-learning
by: Soffair, Nitsan, et al.
Published: (2024)
by: Soffair, Nitsan, et al.
Published: (2024)
Expert Q-learning: Deep Reinforcement Learning with Coarse State Values from Offline Expert Examples
by: Meng, Li, et al.
Published: (2021)
by: Meng, Li, et al.
Published: (2021)
Interactive Double Deep Q-network: Integrating Human Interventions and Evaluative Predictions in Reinforcement Learning of Autonomous Driving
by: Sygkounas, Alkis, et al.
Published: (2025)
by: Sygkounas, Alkis, et al.
Published: (2025)
Forager: a lightweight testbed for continual learning with partial observability in RL
by: Tang, Steven, et al.
Published: (2026)
by: Tang, Steven, et al.
Published: (2026)
Investigating the Interplay of Prioritized Replay and Generalization
by: Panahi, Parham Mohammad, et al.
Published: (2024)
by: Panahi, Parham Mohammad, et al.
Published: (2024)
Trajectory-Aware Eligibility Traces for Off-Policy Reinforcement Learning
by: Daley, Brett, et al.
Published: (2023)
by: Daley, Brett, et al.
Published: (2023)
Stabilizing Extreme Q-learning by Maclaurin Expansion
by: Omura, Motoki, et al.
Published: (2024)
by: Omura, Motoki, et al.
Published: (2024)
Q-learning with Adjoint Matching
by: Li, Qiyang, et al.
Published: (2026)
by: Li, Qiyang, et al.
Published: (2026)
Value Bonuses using Ensemble Errors for Exploration in Reinforcement Learning
by: Wahab, Abdul, et al.
Published: (2026)
by: Wahab, Abdul, et al.
Published: (2026)
Revisiting Mixture Policies in Entropy-Regularized Actor-Critic
by: He, Jiamin, et al.
Published: (2026)
by: He, Jiamin, et al.
Published: (2026)
Universal Approximation Theorem of Deep Q-Networks
by: Qi, Qian
Published: (2025)
by: Qi, Qian
Published: (2025)
Yes, Q-learning Helps Offline In-Context RL
by: Tarasov, Denis, et al.
Published: (2025)
by: Tarasov, Denis, et al.
Published: (2025)
Exclusively Penalized Q-learning for Offline Reinforcement Learning
by: Yeom, Junghyuk, et al.
Published: (2024)
by: Yeom, Junghyuk, et al.
Published: (2024)
On The Presence of Double-Descent in Deep Reinforcement Learning
by: Veselý, Viktor, et al.
Published: (2025)
by: Veselý, Viktor, et al.
Published: (2025)
Exploiting Estimation Bias in Clipped Double Q-Learning for Continous Control Reinforcement Learning Tasks
by: Turcato, Niccolò, et al.
Published: (2024)
by: Turcato, Niccolò, et al.
Published: (2024)
Distributions as Actions: A Unified Framework for Diverse Action Spaces
by: He, Jiamin, et al.
Published: (2025)
by: He, Jiamin, et al.
Published: (2025)
Deep sequence models tend to memorize geometrically; it is unclear why
by: Noroozizadeh, Shahriar, et al.
Published: (2025)
by: Noroozizadeh, Shahriar, et al.
Published: (2025)
UNIQ: Offline Inverse Q-learning for Avoiding Undesirable Demonstrations
by: Hoang, Huy, et al.
Published: (2024)
by: Hoang, Huy, et al.
Published: (2024)
A New View on Planning in Online Reinforcement Learning
by: Roice, Kevin, et al.
Published: (2024)
by: Roice, Kevin, et al.
Published: (2024)
Time After Time: Deep-Q Effect Estimation for Interventions on When and What to do
by: Wald, Yoav, et al.
Published: (2025)
by: Wald, Yoav, et al.
Published: (2025)
$β$-DQN: Improving Deep Q-Learning By Evolving the Behavior
by: Zhang, Hongming, et al.
Published: (2025)
by: Zhang, Hongming, et al.
Published: (2025)
Similar Items
-
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values
by: Daley, Brett, et al.
Published: (2025) -
Demystifying the Recency Heuristic in Temporal-Difference Learning
by: Daley, Brett, et al.
Published: (2024) -
Deep Reinforcement Learning with Gradient Eligibility Traces
by: Elelimy, Esraa, et al.
Published: (2025) -
When is Offline Policy Selection Sample Efficient for Reinforcement Learning?
by: Liu, Vincent, et al.
Published: (2023) -
Harnessing Discrete Representations For Continual Reinforcement Learning
by: Meyer, Edan, et al.
Published: (2023)