Stabilizing Extreme Q-learning by Maclaurin Expansion
Fuente:
arXiv
Saved in:
| Main Authors: | Omura, Motoki, Osa, Takayuki, Mukuta, Yusuke, Harada, Tatsuya |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning
by: Omura, Motoki, et al.
Published: (2024)
by: Omura, Motoki, et al.
Published: (2024)
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning
by: Omura, Motoki, et al.
Published: (2025)
by: Omura, Motoki, et al.
Published: (2025)
Offline Reinforcement Learning with Wasserstein Regularization via Optimal Transport Maps
by: Omura, Motoki, et al.
Published: (2025)
by: Omura, Motoki, et al.
Published: (2025)
Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player Games
by: Ota, Kazuki, et al.
Published: (2026)
by: Ota, Kazuki, et al.
Published: (2026)
Rethinking Policy Diversity in Ensemble Policy Gradient in Large-Scale Reinforcement Learning
by: Shitanda, Naoki, et al.
Published: (2026)
by: Shitanda, Naoki, et al.
Published: (2026)
Cross-Embodiment Offline Reinforcement Learning for Heterogeneous Robot Datasets
by: Abe, Haruki, et al.
Published: (2026)
by: Abe, Haruki, et al.
Published: (2026)
Discovering Multiple Solutions from a Single Task in Offline Reinforcement Learning
by: Osa, Takayuki, et al.
Published: (2024)
by: Osa, Takayuki, et al.
Published: (2024)
Robustifying a Policy in Multi-Agent RL with Diverse Cooperative Behaviors and Adversarial Style Sampling for Assistive Tasks
by: Osa, Takayuki, et al.
Published: (2024)
by: Osa, Takayuki, et al.
Published: (2024)
HyperVQ: MLR-based Vector Quantization in Hyperbolic Space
by: Goswami, Nabarun, et al.
Published: (2024)
by: Goswami, Nabarun, et al.
Published: (2024)
Offline Reinforcement Learning from Datasets with Structured Non-Stationarity
by: Ackermann, Johannes, et al.
Published: (2024)
by: Ackermann, Johannes, et al.
Published: (2024)
Entropy Controllable Direct Preference Optimization
by: Omura, Motoki, et al.
Published: (2024)
by: Omura, Motoki, et al.
Published: (2024)
Deep Double Q-learning
by: Nagarajan, Prabhat, et al.
Published: (2025)
by: Nagarajan, Prabhat, et al.
Published: (2025)
A* Search Without Expansions: Learning Heuristic Functions with Deep Q-Networks
by: Agostinelli, Forest, et al.
Published: (2021)
by: Agostinelli, Forest, et al.
Published: (2021)
MinMaxMin $Q$-learning
by: Soffair, Nitsan, et al.
Published: (2024)
by: Soffair, Nitsan, et al.
Published: (2024)
Is Q-learning an Ill-posed Problem?
by: Wissmann, Philipp, et al.
Published: (2025)
by: Wissmann, Philipp, et al.
Published: (2025)
Q-learning with Adjoint Matching
by: Li, Qiyang, et al.
Published: (2026)
by: Li, Qiyang, et al.
Published: (2026)
Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic
by: Lee, Jeong Woon, et al.
Published: (2026)
by: Lee, Jeong Woon, et al.
Published: (2026)
Exclusively Penalized Q-learning for Offline Reinforcement Learning
by: Yeom, Junghyuk, et al.
Published: (2024)
by: Yeom, Junghyuk, et al.
Published: (2024)
Yes, Q-learning Helps Offline In-Context RL
by: Tarasov, Denis, et al.
Published: (2025)
by: Tarasov, Denis, et al.
Published: (2025)
R2-Dreamer: Redundancy-Reduced World Models without Decoders or Augmentation
by: Morihira, Naoki, et al.
Published: (2026)
by: Morihira, Naoki, et al.
Published: (2026)
UNIQ: Offline Inverse Q-learning for Avoiding Undesirable Demonstrations
by: Hoang, Huy, et al.
Published: (2024)
by: Hoang, Huy, et al.
Published: (2024)
High Precision Audience Expansion via Extreme Classification in a Two-Sided Marketplace
by: Davis, Dillon, et al.
Published: (2026)
by: Davis, Dillon, et al.
Published: (2026)
Compositional Q-learning for electrolyte repletion with imbalanced patient sub-populations
by: Mandyam, Aishwarya, et al.
Published: (2021)
by: Mandyam, Aishwarya, et al.
Published: (2021)
Extreme value forecasting using relevance-based data augmentation with deep learning models
by: Hua, Junru, et al.
Published: (2025)
by: Hua, Junru, et al.
Published: (2025)
SPEQ: Offline Stabilization Phases for Efficient Q-Learning in High Update-To-Data Ratio Reinforcement Learning
by: Romeo, Carlo, et al.
Published: (2025)
by: Romeo, Carlo, et al.
Published: (2025)
Generating In-store Customer Journeys from Scratch with GPT Architectures
by: Horikomi, Taizo, et al.
Published: (2024)
by: Horikomi, Taizo, et al.
Published: (2024)
Load-Aware Training Scheduling for Model Circulation-based Decentralized Federated Learning
by: Kainuma, Haruki, et al.
Published: (2025)
by: Kainuma, Haruki, et al.
Published: (2025)
Understanding the theoretical properties of projected Bellman equation, linear Q-learning, and approximate value iteration
by: Lim, Han-Dong, et al.
Published: (2025)
by: Lim, Han-Dong, et al.
Published: (2025)
DIAR: Diffusion-model-guided Implicit Q-learning with Adaptive Revaluation
by: Park, Jaehyun, et al.
Published: (2024)
by: Park, Jaehyun, et al.
Published: (2024)
ExtremeCast: Boosting Extreme Value Prediction for Global Weather Forecast
by: Xu, Wanghan, et al.
Published: (2024)
by: Xu, Wanghan, et al.
Published: (2024)
UniExtreme: A Universal Foundation Model for Extreme Weather Forecasting
by: Ni, Hang, et al.
Published: (2025)
by: Ni, Hang, et al.
Published: (2025)
Expert Q-learning: Deep Reinforcement Learning with Coarse State Values from Offline Expert Examples
by: Meng, Li, et al.
Published: (2021)
by: Meng, Li, et al.
Published: (2021)
Stochastic Q-learning for Large Discrete Action Spaces
by: Fourati, Fares, et al.
Published: (2024)
by: Fourati, Fares, et al.
Published: (2024)
A finite time analysis of distributed Q-learning
by: Lim, Han-Dong, et al.
Published: (2024)
by: Lim, Han-Dong, et al.
Published: (2024)
Policy Gradient with Tree Expansion
by: Dalal, Gal, et al.
Published: (2023)
by: Dalal, Gal, et al.
Published: (2023)
Probing Implicit Bias in Semi-gradient Q-learning: Visualizing the Effective Loss Landscapes via the Fokker--Planck Equation
by: Yin, Shuyu, et al.
Published: (2024)
by: Yin, Shuyu, et al.
Published: (2024)
Extreme Region Policy Distillation
by: Chen, Changyu, et al.
Published: (2026)
by: Chen, Changyu, et al.
Published: (2026)
Deep Reinforcement Learning with Spiking Q-learning
by: Chen, Ding, et al.
Published: (2022)
by: Chen, Ding, et al.
Published: (2022)
Tensor tree learns hidden relational structures in data to construct generative models
by: Harada, Kenji, et al.
Published: (2024)
by: Harada, Kenji, et al.
Published: (2024)
Transformation Categorization Based on Group Decomposition Theory Using Parameter Division
by: Komatsu, Takayuki, et al.
Published: (2026)
by: Komatsu, Takayuki, et al.
Published: (2026)
Similar Items
-
Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning
by: Omura, Motoki, et al.
Published: (2024) -
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning
by: Omura, Motoki, et al.
Published: (2025) -
Offline Reinforcement Learning with Wasserstein Regularization via Optimal Transport Maps
by: Omura, Motoki, et al.
Published: (2025) -
Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player Games
by: Ota, Kazuki, et al.
Published: (2026) -
Rethinking Policy Diversity in Ensemble Policy Gradient in Large-Scale Reinforcement Learning
by: Shitanda, Naoki, et al.
Published: (2026)