Implicit Q-Learning and SARSA: Liberating Policy Control from Step-Size Calibration
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Hwanwoo, Laber, Eric |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stabilizing Temporal Difference Learning via Implicit Stochastic Recursion
by: Kim, Hwanwoo, et al.
Published: (2025)
by: Kim, Hwanwoo, et al.
Published: (2025)
Implicit Updates for Average-Reward Temporal Difference Learning
by: Kim, Hwanwoo, et al.
Published: (2025)
by: Kim, Hwanwoo, et al.
Published: (2025)
Adaptive Policy Learning Under Unknown Network Interference
by: Gleich, Aidan, et al.
Published: (2026)
by: Gleich, Aidan, et al.
Published: (2026)
Convergence Guarantees for Federated SARSA with Local Training and Heterogeneous Agents
by: Mangold, Paul, et al.
Published: (2025)
by: Mangold, Paul, et al.
Published: (2025)
Scalable Policy Maximization Under Network Interference
by: Gleich, Aidan, et al.
Published: (2025)
by: Gleich, Aidan, et al.
Published: (2025)
Enhancing Gaussian Process Surrogates for Optimization and Posterior Approximation via Random Exploration
by: Kim, Hwanwoo, et al.
Published: (2024)
by: Kim, Hwanwoo, et al.
Published: (2024)
Bayesian Optimization with Inexact Acquisition: Is Random Grid Search Sufficient?
by: Kim, Hwanwoo, et al.
Published: (2025)
by: Kim, Hwanwoo, et al.
Published: (2025)
Exploiting Concavity Information in Gaussian Process Contextual Bandit Optimization
by: Li, Kevin, et al.
Published: (2025)
by: Li, Kevin, et al.
Published: (2025)
Optimization of Inter-group Criteria for Clustering with Minimum Size Constraints
by: Laber, Eduardo S., et al.
Published: (2024)
by: Laber, Eduardo S., et al.
Published: (2024)
Empirical Bound Information-Directed Sampling for Norm-Agnostic Bandits
by: Suder, Piotr M., et al.
Published: (2025)
by: Suder, Piotr M., et al.
Published: (2025)
ReTaSA: A Nonparametric Functional Estimation Approach for Addressing Continuous Target Shift
by: Kim, Hwanwoo, et al.
Published: (2024)
by: Kim, Hwanwoo, et al.
Published: (2024)
AlignIQL: Policy Alignment in Implicit Q-Learning through Constrained Optimization
by: He, Longxiang, et al.
Published: (2024)
by: He, Longxiang, et al.
Published: (2024)
State-Separated SARSA: A Practical Sequential Decision-Making Algorithm with Recovering Rewards
by: Tanimoto, Yuto, et al.
Published: (2024)
by: Tanimoto, Yuto, et al.
Published: (2024)
Enhancing Policy Gradient with the Polyak Step-Size Adaption
by: Li, Yunxiang, et al.
Published: (2024)
by: Li, Yunxiang, et al.
Published: (2024)
Two-Step Q-Learning
by: Vijesh, Antony, et al.
Published: (2024)
by: Vijesh, Antony, et al.
Published: (2024)
Efficient optimization of expensive black-box simulators via marginal means, with application to neutrino detector design
by: Kim, Hwanwoo, et al.
Published: (2025)
by: Kim, Hwanwoo, et al.
Published: (2025)
One-Step Flow Q-Learning: Addressing the Diffusion Policy Bottleneck in Offline Reinforcement Learning
by: Nguyen, Thanh, et al.
Published: (2025)
by: Nguyen, Thanh, et al.
Published: (2025)
One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
by: Wang, Zeyuan, et al.
Published: (2025)
by: Wang, Zeyuan, et al.
Published: (2025)
Semi-Gradient SARSA Routing with Theoretical Guarantee on Traffic Stability and Weight Convergence
by: Wu, Yidan, et al.
Published: (2025)
by: Wu, Yidan, et al.
Published: (2025)
IQL-TD-MPC: Implicit Q-Learning for Hierarchical Model Predictive Control
by: Chitnis, Rohan, et al.
Published: (2023)
by: Chitnis, Rohan, et al.
Published: (2023)
New bounds on the cohesion of complete-link and other linkage methods for agglomeration clustering
by: Dasgupta, Sanjoy, et al.
Published: (2024)
by: Dasgupta, Sanjoy, et al.
Published: (2024)
DIAR: Diffusion-model-guided Implicit Q-learning with Adaptive Revaluation
by: Park, Jaehyun, et al.
Published: (2024)
by: Park, Jaehyun, et al.
Published: (2024)
Segmenting Action-Value Functions Over Time-Scales in SARSA via TD($Δ$)
by: Humayoo, Mahammad
Published: (2024)
by: Humayoo, Mahammad
Published: (2024)
Sequential Knockoffs for Variable Selection in Reinforcement Learning
by: Ma, Tao, et al.
Published: (2023)
by: Ma, Tao, et al.
Published: (2023)
On Calibration in Multi-Distribution Learning
by: Verma, Rajeev, et al.
Published: (2024)
by: Verma, Rajeev, et al.
Published: (2024)
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
by: Deng, Yuntian, et al.
Published: (2024)
by: Deng, Yuntian, et al.
Published: (2024)
Efficient Hierarchical Implicit Flow Q-learning for Offline Goal-conditioned Reinforcement Learning
by: Dong, Zhiqiang, et al.
Published: (2026)
by: Dong, Zhiqiang, et al.
Published: (2026)
PIQL: Projective Implicit Q-Learning with Support Constraint for Offline Reinforcement Learning
by: Han, Xinchen, et al.
Published: (2025)
by: Han, Xinchen, et al.
Published: (2025)
Navigation with QPHIL: Quantizing Planner for Hierarchical Implicit Q-Learning
by: Canesse, Alexi, et al.
Published: (2024)
by: Canesse, Alexi, et al.
Published: (2024)
CafeQ: Calibration-free Quantization via Learned Transformations and Adaptive Rounding
by: Sun, Ziteng, et al.
Published: (2025)
by: Sun, Ziteng, et al.
Published: (2025)
Step-Size Decay and Structural Stagnation in Greedy Sparse Learning
by: Berná, Pablo M.
Published: (2026)
by: Berná, Pablo M.
Published: (2026)
Exploring Time-Step Size in Reinforcement Learning for Sepsis Treatment
by: Sun, Yingchuan, et al.
Published: (2025)
by: Sun, Yingchuan, et al.
Published: (2025)
Faster Game Solving via Asymmetry of Step Sizes
by: Meng, Linjian, et al.
Published: (2025)
by: Meng, Linjian, et al.
Published: (2025)
Digi-Q: Learning Q-Value Functions for Training Device-Control Agents
by: Bai, Hao, et al.
Published: (2025)
by: Bai, Hao, et al.
Published: (2025)
Truncated Rectified Flow Policy for Reinforcement Learning with One-Step Sampling
by: Zhou, Xubin, et al.
Published: (2026)
by: Zhou, Xubin, et al.
Published: (2026)
Active Learning with Selective Time-Step Acquisition for PDEs
by: Kim, Yegon, et al.
Published: (2025)
by: Kim, Yegon, et al.
Published: (2025)
An Efficient On-Policy Deep Learning Framework for Stochastic Optimal Control
by: Hua, Mengjian, et al.
Published: (2024)
by: Hua, Mengjian, et al.
Published: (2024)
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning
by: Vu, Minh, et al.
Published: (2025)
by: Vu, Minh, et al.
Published: (2025)
Frictional Q-Learning
by: Kim, Hyunwoo, et al.
Published: (2025)
by: Kim, Hyunwoo, et al.
Published: (2025)
Adaptive $Q$-Aid for Conditional Supervised Learning in Offline Reinforcement Learning
by: Kim, Jeonghye, et al.
Published: (2024)
by: Kim, Jeonghye, et al.
Published: (2024)
Similar Items
-
Stabilizing Temporal Difference Learning via Implicit Stochastic Recursion
by: Kim, Hwanwoo, et al.
Published: (2025) -
Implicit Updates for Average-Reward Temporal Difference Learning
by: Kim, Hwanwoo, et al.
Published: (2025) -
Adaptive Policy Learning Under Unknown Network Interference
by: Gleich, Aidan, et al.
Published: (2026) -
Convergence Guarantees for Federated SARSA with Local Training and Heterogeneous Agents
by: Mangold, Paul, et al.
Published: (2025) -
Scalable Policy Maximization Under Network Interference
by: Gleich, Aidan, et al.
Published: (2025)