Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Omura, Motoki, Ota, Kazuki, Osa, Takayuki, Mukuta, Yusuke, Harada, Tatsuya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning
by: Omura, Motoki, et al.
Published: (2024)
by: Omura, Motoki, et al.
Published: (2024)
Offline Reinforcement Learning with Wasserstein Regularization via Optimal Transport Maps
by: Omura, Motoki, et al.
Published: (2025)
by: Omura, Motoki, et al.
Published: (2025)
Stabilizing Extreme Q-learning by Maclaurin Expansion
by: Omura, Motoki, et al.
Published: (2024)
by: Omura, Motoki, et al.
Published: (2024)
Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player Games
by: Ota, Kazuki, et al.
Published: (2026)
by: Ota, Kazuki, et al.
Published: (2026)
Rethinking Policy Diversity in Ensemble Policy Gradient in Large-Scale Reinforcement Learning
by: Shitanda, Naoki, et al.
Published: (2026)
by: Shitanda, Naoki, et al.
Published: (2026)
Cross-Embodiment Offline Reinforcement Learning for Heterogeneous Robot Datasets
by: Abe, Haruki, et al.
Published: (2026)
by: Abe, Haruki, et al.
Published: (2026)
Robustifying a Policy in Multi-Agent RL with Diverse Cooperative Behaviors and Adversarial Style Sampling for Assistive Tasks
by: Osa, Takayuki, et al.
Published: (2024)
by: Osa, Takayuki, et al.
Published: (2024)
Parameterized Projected Bellman Operator
by: Vincent, Théo, et al.
Published: (2023)
by: Vincent, Théo, et al.
Published: (2023)
MICRO: Model-Based Offline Reinforcement Learning with a Conservative Bellman Operator
by: Liu, Xiao-Yin, et al.
Published: (2023)
by: Liu, Xiao-Yin, et al.
Published: (2023)
Discovering Multiple Solutions from a Single Task in Offline Reinforcement Learning
by: Osa, Takayuki, et al.
Published: (2024)
by: Osa, Takayuki, et al.
Published: (2024)
Theoretical Barriers in Bellman-Based Reinforcement Learning
by: Pinon, Brieuc, et al.
Published: (2025)
by: Pinon, Brieuc, et al.
Published: (2025)
Linear Bellman Completeness Suffices for Efficient Online Reinforcement Learning with Few Actions
by: Golowich, Noah, et al.
Published: (2024)
by: Golowich, Noah, et al.
Published: (2024)
Bellman Error Centering
by: Chen, Xingguo, et al.
Published: (2025)
by: Chen, Xingguo, et al.
Published: (2025)
Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity
by: Muni, Aneri, et al.
Published: (2026)
by: Muni, Aneri, et al.
Published: (2026)
Bellman Diffusion Models
by: Schramm, Liam, et al.
Published: (2024)
by: Schramm, Liam, et al.
Published: (2024)
Continuous Reasoning for Vision-Language-Action
by: Wu, Yueh-Hua, et al.
Published: (2026)
by: Wu, Yueh-Hua, et al.
Published: (2026)
Offline Reinforcement Learning from Datasets with Structured Non-Stationarity
by: Ackermann, Johannes, et al.
Published: (2024)
by: Ackermann, Johannes, et al.
Published: (2024)
Path-Coupled Bellman Flows for Distributional Reinforcement Learning
by: Xu, Boyang, et al.
Published: (2026)
by: Xu, Boyang, et al.
Published: (2026)
The Role of Inherent Bellman Error in Offline Reinforcement Learning with Linear Function Approximation
by: Golowich, Noah, et al.
Published: (2024)
by: Golowich, Noah, et al.
Published: (2024)
Tactical Decision Making for Autonomous Trucks by Deep Reinforcement Learning with Total Cost of Operation Based Reward
by: Pathare, Deepthi, et al.
Published: (2024)
by: Pathare, Deepthi, et al.
Published: (2024)
On the Uniqueness of Solution for the Bellman Equation of LTL Objectives
by: Xuan, Zetong, et al.
Published: (2024)
by: Xuan, Zetong, et al.
Published: (2024)
Iterated $Q$-Network: Beyond One-Step Bellman Updates in Deep Reinforcement Learning
by: Vincent, Théo, et al.
Published: (2024)
by: Vincent, Théo, et al.
Published: (2024)
A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
by: Patterson, Andrew, et al.
Published: (2021)
by: Patterson, Andrew, et al.
Published: (2021)
Prospect-Theory Behavior from Bellman Optimality in MDPs with Catastrophic States
by: Chen, Yujiao
Published: (2026)
by: Chen, Yujiao
Published: (2026)
Bellman Unbiasedness: Toward Provably Efficient Distributional Reinforcement Learning with General Value Function Approximation
by: Cho, Taehyun, et al.
Published: (2024)
by: Cho, Taehyun, et al.
Published: (2024)
Online Training and Pruning of Deep Reinforcement Learning Networks
by: Guenter, Valentin Frank Ingmar, et al.
Published: (2025)
by: Guenter, Valentin Frank Ingmar, et al.
Published: (2025)
R2-Dreamer: Redundancy-Reduced World Models without Decoders or Augmentation
by: Morihira, Naoki, et al.
Published: (2026)
by: Morihira, Naoki, et al.
Published: (2026)
ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles
by: Zhao, Kai, et al.
Published: (2023)
by: Zhao, Kai, et al.
Published: (2023)
Bellman operator convergence enhancements in reinforcement learning algorithms
by: Kadurha, David Krame, et al.
Published: (2025)
by: Kadurha, David Krame, et al.
Published: (2025)
Leveraging Sub-Optimal Data for Human-in-the-Loop Reinforcement Learning
by: Muslimani, Calarina, et al.
Published: (2024)
by: Muslimani, Calarina, et al.
Published: (2024)
A New Perspective on Transformers in Online Reinforcement Learning for Continuous Control
by: Kachaev, Nikita, et al.
Published: (2025)
by: Kachaev, Nikita, et al.
Published: (2025)
Accelerated Online Reinforcement Learning using Auxiliary Start State Distributions
by: Mehra, Aman, et al.
Published: (2025)
by: Mehra, Aman, et al.
Published: (2025)
Robot Policy Transfer with Online Demonstrations: An Active Reinforcement Learning Approach
by: Hou, Muhan, et al.
Published: (2025)
by: Hou, Muhan, et al.
Published: (2025)
Temporal Distance-aware Transition Augmentation for Offline Model-based Reinforcement Learning
by: Lee, Dongsu, et al.
Published: (2025)
by: Lee, Dongsu, et al.
Published: (2025)
Sample Complexity of Distributionally Robust Off-Dynamics Reinforcement Learning with Online Interaction
by: He, Yiting, et al.
Published: (2025)
by: He, Yiting, et al.
Published: (2025)
A Plug-and-Play Fully On-the-Job Real-Time Reinforcement Learning Algorithm for a Direct-Drive Tandem-Wing Experiment Platforms Under Multiple Random Operating Conditions
by: Minghao, Zhang, et al.
Published: (2024)
by: Minghao, Zhang, et al.
Published: (2024)
Visualizing Critic Match Loss Landscapes for Interpretation of Online Reinforcement Learning Control Algorithms
by: Liu, Jingyi, et al.
Published: (2026)
by: Liu, Jingyi, et al.
Published: (2026)
Adaptive Regularization of Representation Rank as an Implicit Constraint of Bellman Equation
by: He, Qiang, et al.
Published: (2024)
by: He, Qiang, et al.
Published: (2024)
Failure-Aware RL: Reliable Offline-to-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation
by: Li, Huanyu, et al.
Published: (2026)
by: Li, Huanyu, et al.
Published: (2026)
GOPT: Generalizable Online 3D Bin Packing via Transformer-based Deep Reinforcement Learning
by: Xiong, Heng, et al.
Published: (2024)
by: Xiong, Heng, et al.
Published: (2024)
Similar Items
-
Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning
by: Omura, Motoki, et al.
Published: (2024) -
Offline Reinforcement Learning with Wasserstein Regularization via Optimal Transport Maps
by: Omura, Motoki, et al.
Published: (2025) -
Stabilizing Extreme Q-learning by Maclaurin Expansion
by: Omura, Motoki, et al.
Published: (2024) -
Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player Games
by: Ota, Kazuki, et al.
Published: (2026) -
Rethinking Policy Diversity in Ensemble Policy Gradient in Large-Scale Reinforcement Learning
by: Shitanda, Naoki, et al.
Published: (2026)