Implicit Updates for Average-Reward Temporal Difference Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Hwanwoo, Cho, Dongkyu Derek, Laber, Eric |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stabilizing Temporal Difference Learning via Implicit Stochastic Recursion
by: Kim, Hwanwoo, et al.
Published: (2025)
by: Kim, Hwanwoo, et al.
Published: (2025)
Implicit Q-Learning and SARSA: Liberating Policy Control from Step-Size Calibration
by: Kim, Hwanwoo, et al.
Published: (2026)
by: Kim, Hwanwoo, et al.
Published: (2026)
Enhancing Gaussian Process Surrogates for Optimization and Posterior Approximation via Random Exploration
by: Kim, Hwanwoo, et al.
Published: (2024)
by: Kim, Hwanwoo, et al.
Published: (2024)
Dealing with the Evil Twins: Improving Random Augmentation by Addressing Catastrophic Forgetting of Diverse Augmentations
by: Cho, Dongkyu, et al.
Published: (2025)
by: Cho, Dongkyu, et al.
Published: (2025)
Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes
by: Blaser, Ethan, et al.
Published: (2026)
by: Blaser, Ethan, et al.
Published: (2026)
Bayesian Optimization with Inexact Acquisition: Is Random Grid Search Sufficient?
by: Kim, Hwanwoo, et al.
Published: (2025)
by: Kim, Hwanwoo, et al.
Published: (2025)
Exploiting Concavity Information in Gaussian Process Contextual Bandit Optimization
by: Li, Kevin, et al.
Published: (2025)
by: Li, Kevin, et al.
Published: (2025)
Adaptive Policy Learning Under Unknown Network Interference
by: Gleich, Aidan, et al.
Published: (2026)
by: Gleich, Aidan, et al.
Published: (2026)
Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs
by: Cho, Dongkyu Derek, et al.
Published: (2025)
by: Cho, Dongkyu Derek, et al.
Published: (2025)
Empirical Bound Information-Directed Sampling for Norm-Agnostic Bandits
by: Suder, Piotr M., et al.
Published: (2025)
by: Suder, Piotr M., et al.
Published: (2025)
ReTaSA: A Nonparametric Functional Estimation Approach for Addressing Continuous Target Shift
by: Kim, Hwanwoo, et al.
Published: (2024)
by: Kim, Hwanwoo, et al.
Published: (2024)
Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration
by: Cho, Dongkyu, et al.
Published: (2025)
by: Cho, Dongkyu, et al.
Published: (2025)
Scalable Policy Maximization Under Network Interference
by: Gleich, Aidan, et al.
Published: (2025)
by: Gleich, Aidan, et al.
Published: (2025)
Forget Forgetting: Continual Learning in a World of Abundant Memory
by: Cho, Dongkyu, et al.
Published: (2025)
by: Cho, Dongkyu, et al.
Published: (2025)
Learning to Explore in Diverse Reward Settings via Temporal-Difference-Error Maximization
by: Griesbach, Sebastian, et al.
Published: (2025)
by: Griesbach, Sebastian, et al.
Published: (2025)
Efficient optimization of expensive black-box simulators via marginal means, with application to neutrino detector design
by: Kim, Hwanwoo, et al.
Published: (2025)
by: Kim, Hwanwoo, et al.
Published: (2025)
Tree of Concepts: Interpretable Continual Learners in Non-Stationary Clinical Domains
by: Cho, Dongkyu, et al.
Published: (2026)
by: Cho, Dongkyu, et al.
Published: (2026)
Average-Reward Maximum Entropy Reinforcement Learning for Underactuated Double Pendulum Tasks
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
Reward-Zero: Language Embedding Driven Implicit Reward Mechanisms for Reinforcement Learning
by: Zhang, Heng, et al.
Published: (2026)
by: Zhang, Heng, et al.
Published: (2026)
Generalized Gaussian Temporal Difference Error for Uncertainty-aware Reinforcement Learning
by: Kim, Seyeon, et al.
Published: (2024)
by: Kim, Seyeon, et al.
Published: (2024)
Provably Sample-Efficient Robust Reinforcement Learning with Average Reward
by: Roch, Zachary, et al.
Published: (2025)
by: Roch, Zachary, et al.
Published: (2025)
Bandit Simulation for Average Reward Inference
by: Praharaj, Samya, et al.
Published: (2026)
by: Praharaj, Samya, et al.
Published: (2026)
NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models
by: Chong, Hyochan, et al.
Published: (2026)
by: Chong, Hyochan, et al.
Published: (2026)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
by: Hong, Kihyuk, et al.
Published: (2024)
by: Hong, Kihyuk, et al.
Published: (2024)
Efficient Learning of POMDPs with Known Observation Model in Average-Reward Setting
by: Russo, Alessio, et al.
Published: (2024)
by: Russo, Alessio, et al.
Published: (2024)
Finite-Sample Analysis of Policy Evaluation for Robust Average Reward Reinforcement Learning
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
by: Chen, Zijun, et al.
Published: (2025)
by: Chen, Zijun, et al.
Published: (2025)
Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces
by: Kar, Avik, et al.
Published: (2024)
by: Kar, Avik, et al.
Published: (2024)
Neural-network quantum state study of the long-range antiferromagnetic Ising chain
by: Kim, Jicheol, et al.
Published: (2023)
by: Kim, Jicheol, et al.
Published: (2023)
Temporal Difference Learning with Compressed Updates: Error-Feedback meets Reinforcement Learning
by: Mitra, Aritra, et al.
Published: (2023)
by: Mitra, Aritra, et al.
Published: (2023)
Average-Reward Soft Actor-Critic
by: Adamczyk, Jacob, et al.
Published: (2025)
by: Adamczyk, Jacob, et al.
Published: (2025)
Sample Complexity of Average-Reward Q-Learning: From Single-agent to Federated Reinforcement Learning
by: Jiao, Yuchen, et al.
Published: (2026)
by: Jiao, Yuchen, et al.
Published: (2026)
Optimization of Inter-group Criteria for Clustering with Minimum Size Constraints
by: Laber, Eduardo S., et al.
Published: (2024)
by: Laber, Eduardo S., et al.
Published: (2024)
New bounds on the cohesion of complete-link and other linkage methods for agglomeration clustering
by: Dasgupta, Sanjoy, et al.
Published: (2024)
by: Dasgupta, Sanjoy, et al.
Published: (2024)
Sparse Autoencoders Reveal Temporal Difference Learning in Large Language Models
by: Demircan, Can, et al.
Published: (2024)
by: Demircan, Can, et al.
Published: (2024)
TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference
by: Zhang, Dan, et al.
Published: (2025)
by: Zhang, Dan, et al.
Published: (2025)
Uncertainty-Aware Multi-Objective Reinforcement Learning-Guided Diffusion Models for 3D De Novo Molecular Design
by: Chen, Lianghong, et al.
Published: (2025)
by: Chen, Lianghong, et al.
Published: (2025)
Policy Zooming: Adaptive Discretization-based Infinite-Horizon Average-Reward Reinforcement Learning
by: Kar, Avik, et al.
Published: (2024)
by: Kar, Avik, et al.
Published: (2024)
Personalized Multi-Agent Average Reward TD-Learning via Joint Linear Approximation
by: Wang, Leo Muxing, et al.
Published: (2026)
by: Wang, Leo Muxing, et al.
Published: (2026)
A Harmonic Mean Formulation of Average Reward Reinforcement Learning in SMDPs
by: Shtossel, Erel, et al.
Published: (2026)
by: Shtossel, Erel, et al.
Published: (2026)
Similar Items
-
Stabilizing Temporal Difference Learning via Implicit Stochastic Recursion
by: Kim, Hwanwoo, et al.
Published: (2025) -
Implicit Q-Learning and SARSA: Liberating Policy Control from Step-Size Calibration
by: Kim, Hwanwoo, et al.
Published: (2026) -
Enhancing Gaussian Process Surrogates for Optimization and Posterior Approximation via Random Exploration
by: Kim, Hwanwoo, et al.
Published: (2024) -
Dealing with the Evil Twins: Improving Random Augmentation by Addressing Catastrophic Forgetting of Diverse Augmentations
by: Cho, Dongkyu, et al.
Published: (2025) -
Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes
by: Blaser, Ethan, et al.
Published: (2026)