Iterated $Q$-Network: Beyond One-Step Bellman Updates in Deep Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Vincent, Théo, Palenicek, Daniel, Belousov, Boris, Peters, Jan, D'Eramo, Carlo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Adaptive $Q$-Network: On-the-fly Target Selection for Deep Reinforcement Learning
por: Vincent, Théo, et al.
Publicado: (2024)
por: Vincent, Théo, et al.
Publicado: (2024)
Parameterized Projected Bellman Operator
por: Vincent, Théo, et al.
Publicado: (2023)
por: Vincent, Théo, et al.
Publicado: (2023)
Eau De $Q$-Network: Adaptive Distillation of Neural Networks in Deep Reinforcement Learning
por: Vincent, Théo, et al.
Publicado: (2025)
por: Vincent, Théo, et al.
Publicado: (2025)
$K$-Level Policy Gradients for Multi-Agent Reinforcement Learning
por: Reddi, Aryaman, et al.
Publicado: (2025)
por: Reddi, Aryaman, et al.
Publicado: (2025)
Gradient Iterated Temporal-Difference Learning
por: Vincent, Théo, et al.
Publicado: (2026)
por: Vincent, Théo, et al.
Publicado: (2026)
Bridging the Performance Gap Between Target-Free and Target-Based Reinforcement Learning
por: Vincent, Théo, et al.
Publicado: (2025)
por: Vincent, Théo, et al.
Publicado: (2025)
Scaling CrossQ with Weight Normalization
por: Palenicek, Daniel, et al.
Publicado: (2025)
por: Palenicek, Daniel, et al.
Publicado: (2025)
XQC: Well-conditioned Optimization Accelerates Deep Reinforcement Learning
por: Palenicek, Daniel, et al.
Publicado: (2025)
por: Palenicek, Daniel, et al.
Publicado: (2025)
Deterministic Exploration via Stationary Bellman Error Maximization
por: Griesbach, Sebastian, et al.
Publicado: (2024)
por: Griesbach, Sebastian, et al.
Publicado: (2024)
Multi-Task Reinforcement Learning with Mixture of Orthogonal Experts
por: Hendawy, Ahmed, et al.
Publicado: (2023)
por: Hendawy, Ahmed, et al.
Publicado: (2023)
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning
por: Hendawy, Ahmed, et al.
Publicado: (2025)
por: Hendawy, Ahmed, et al.
Publicado: (2025)
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
por: Palenicek, Daniel, et al.
Publicado: (2025)
por: Palenicek, Daniel, et al.
Publicado: (2025)
CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity
por: Bhatt, Aditya, et al.
Publicado: (2019)
por: Bhatt, Aditya, et al.
Publicado: (2019)
Sharing Knowledge in Multi-Task Deep Reinforcement Learning
por: D'Eramo, Carlo, et al.
Publicado: (2024)
por: D'Eramo, Carlo, et al.
Publicado: (2024)
On the Benefit of Optimal Transport for Curriculum Reinforcement Learning
por: Klink, Pascal, et al.
Publicado: (2023)
por: Klink, Pascal, et al.
Publicado: (2023)
Streaming Reinforcement Learning under Partial Observability with Real-Time Recurrent Learning
por: Farr, Noah, et al.
Publicado: (2026)
por: Farr, Noah, et al.
Publicado: (2026)
Contact Energy Based Hindsight Experience Prioritization
por: Sayar, Erdi, et al.
Publicado: (2023)
por: Sayar, Erdi, et al.
Publicado: (2023)
Do Not Imitate, Reinforce: Iterative Classification via Belief Refinement
por: Kallel, Mahdi, et al.
Publicado: (2026)
por: Kallel, Mahdi, et al.
Publicado: (2026)
MuJoCo MPC for Humanoid Control: Evaluation on HumanoidBench
por: Meser, Moritz, et al.
Publicado: (2024)
por: Meser, Moritz, et al.
Publicado: (2024)
Beyond Single-Step Updates: Reinforcement Learning of Heuristics with Limited-Horizon Search
por: Hadar, Gal, et al.
Publicado: (2025)
por: Hadar, Gal, et al.
Publicado: (2025)
Deep Reinforcement Learning Agents are not even close to Human Intelligence
por: Delfosse, Quentin, et al.
Publicado: (2025)
por: Delfosse, Quentin, et al.
Publicado: (2025)
Learning to Explore in Diverse Reward Settings via Temporal-Difference-Error Maximization
por: Griesbach, Sebastian, et al.
Publicado: (2025)
por: Griesbach, Sebastian, et al.
Publicado: (2025)
SPEQ: Offline Stabilization Phases for Efficient Q-Learning in High Update-To-Data Ratio Reinforcement Learning
por: Romeo, Carlo, et al.
Publicado: (2025)
por: Romeo, Carlo, et al.
Publicado: (2025)
Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning
por: Omura, Motoki, et al.
Publicado: (2024)
por: Omura, Motoki, et al.
Publicado: (2024)
Cosmic-Ray Signatures of Annihilating and Semi-Annihilating Dark Matter via One-Step Cascades
por: D'Eramo, Francesco, et al.
Publicado: (2026)
por: D'Eramo, Francesco, et al.
Publicado: (2026)
Beyond the Bellman Fixed Point: Geometry and Fast Policy Identification in Value Iteration
por: Lee, Donghwan
Publicado: (2026)
por: Lee, Donghwan
Publicado: (2026)
Theoretical Barriers in Bellman-Based Reinforcement Learning
por: Pinon, Brieuc, et al.
Publicado: (2025)
por: Pinon, Brieuc, et al.
Publicado: (2025)
Light to Heavy, Brief to Eternal: An Axion for Every Occasion (in the Early Universe)
por: D'Eramo, Francesco
Publicado: (2026)
por: D'Eramo, Francesco
Publicado: (2026)
Dynamic Obstacle Avoidance with Bounded Rationality Adversarial Reinforcement Learning
por: Holgado-Alvarez, Jose-Luis, et al.
Publicado: (2025)
por: Holgado-Alvarez, Jose-Luis, et al.
Publicado: (2025)
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning
por: Omura, Motoki, et al.
Publicado: (2025)
por: Omura, Motoki, et al.
Publicado: (2025)
Multi-agent Reinforcement Learning with Deep Networks for Diverse Q-Vectors
por: Luo, Zhenglong, et al.
Publicado: (2024)
por: Luo, Zhenglong, et al.
Publicado: (2024)
One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
por: Wang, Zeyuan, et al.
Publicado: (2025)
por: Wang, Zeyuan, et al.
Publicado: (2025)
Dual-Objective Reinforcement Learning with Novel Hamilton-Jacobi-Bellman Formulations
por: Sharpless, William, et al.
Publicado: (2025)
por: Sharpless, William, et al.
Publicado: (2025)
Path-Coupled Bellman Flows for Distributional Reinforcement Learning
por: Xu, Boyang, et al.
Publicado: (2026)
por: Xu, Boyang, et al.
Publicado: (2026)
Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates
por: Diwan, Anish, et al.
Publicado: (2026)
por: Diwan, Anish, et al.
Publicado: (2026)
Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning
por: Ding, Zihan, et al.
Publicado: (2024)
por: Ding, Zihan, et al.
Publicado: (2024)
Linear Bellman Completeness Suffices for Efficient Online Reinforcement Learning with Few Actions
por: Golowich, Noah, et al.
Publicado: (2024)
por: Golowich, Noah, et al.
Publicado: (2024)
MICRO: Model-Based Offline Reinforcement Learning with a Conservative Bellman Operator
por: Liu, Xiao-Yin, et al.
Publicado: (2023)
por: Liu, Xiao-Yin, et al.
Publicado: (2023)
The Role of Inherent Bellman Error in Offline Reinforcement Learning with Linear Function Approximation
por: Golowich, Noah, et al.
Publicado: (2024)
por: Golowich, Noah, et al.
Publicado: (2024)
Beyond ReLU: Chebyshev-DQN for Enhanced Deep Q-Networks
por: Yazdannik, Saman, et al.
Publicado: (2025)
por: Yazdannik, Saman, et al.
Publicado: (2025)
Ejemplares similares
-
Adaptive $Q$-Network: On-the-fly Target Selection for Deep Reinforcement Learning
por: Vincent, Théo, et al.
Publicado: (2024) -
Parameterized Projected Bellman Operator
por: Vincent, Théo, et al.
Publicado: (2023) -
Eau De $Q$-Network: Adaptive Distillation of Neural Networks in Deep Reinforcement Learning
por: Vincent, Théo, et al.
Publicado: (2025) -
$K$-Level Policy Gradients for Multi-Agent Reinforcement Learning
por: Reddi, Aryaman, et al.
Publicado: (2025) -
Gradient Iterated Temporal-Difference Learning
por: Vincent, Théo, et al.
Publicado: (2026)