Implicit Updates for Average-Reward Temporal Difference Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Hwanwoo, Cho, Dongkyu Derek, Laber, Eric |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stabilizing Temporal Difference Learning via Implicit Stochastic Recursion
von: Kim, Hwanwoo, et al.
Veröffentlicht: (2025)
von: Kim, Hwanwoo, et al.
Veröffentlicht: (2025)
Implicit Q-Learning and SARSA: Liberating Policy Control from Step-Size Calibration
von: Kim, Hwanwoo, et al.
Veröffentlicht: (2026)
von: Kim, Hwanwoo, et al.
Veröffentlicht: (2026)
Enhancing Gaussian Process Surrogates for Optimization and Posterior Approximation via Random Exploration
von: Kim, Hwanwoo, et al.
Veröffentlicht: (2024)
von: Kim, Hwanwoo, et al.
Veröffentlicht: (2024)
Dealing with the Evil Twins: Improving Random Augmentation by Addressing Catastrophic Forgetting of Diverse Augmentations
von: Cho, Dongkyu, et al.
Veröffentlicht: (2025)
von: Cho, Dongkyu, et al.
Veröffentlicht: (2025)
Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes
von: Blaser, Ethan, et al.
Veröffentlicht: (2026)
von: Blaser, Ethan, et al.
Veröffentlicht: (2026)
Bayesian Optimization with Inexact Acquisition: Is Random Grid Search Sufficient?
von: Kim, Hwanwoo, et al.
Veröffentlicht: (2025)
von: Kim, Hwanwoo, et al.
Veröffentlicht: (2025)
Exploiting Concavity Information in Gaussian Process Contextual Bandit Optimization
von: Li, Kevin, et al.
Veröffentlicht: (2025)
von: Li, Kevin, et al.
Veröffentlicht: (2025)
Adaptive Policy Learning Under Unknown Network Interference
von: Gleich, Aidan, et al.
Veröffentlicht: (2026)
von: Gleich, Aidan, et al.
Veröffentlicht: (2026)
Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs
von: Cho, Dongkyu Derek, et al.
Veröffentlicht: (2025)
von: Cho, Dongkyu Derek, et al.
Veröffentlicht: (2025)
Empirical Bound Information-Directed Sampling for Norm-Agnostic Bandits
von: Suder, Piotr M., et al.
Veröffentlicht: (2025)
von: Suder, Piotr M., et al.
Veröffentlicht: (2025)
ReTaSA: A Nonparametric Functional Estimation Approach for Addressing Continuous Target Shift
von: Kim, Hwanwoo, et al.
Veröffentlicht: (2024)
von: Kim, Hwanwoo, et al.
Veröffentlicht: (2024)
Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration
von: Cho, Dongkyu, et al.
Veröffentlicht: (2025)
von: Cho, Dongkyu, et al.
Veröffentlicht: (2025)
Scalable Policy Maximization Under Network Interference
von: Gleich, Aidan, et al.
Veröffentlicht: (2025)
von: Gleich, Aidan, et al.
Veröffentlicht: (2025)
Forget Forgetting: Continual Learning in a World of Abundant Memory
von: Cho, Dongkyu, et al.
Veröffentlicht: (2025)
von: Cho, Dongkyu, et al.
Veröffentlicht: (2025)
Learning to Explore in Diverse Reward Settings via Temporal-Difference-Error Maximization
von: Griesbach, Sebastian, et al.
Veröffentlicht: (2025)
von: Griesbach, Sebastian, et al.
Veröffentlicht: (2025)
Efficient optimization of expensive black-box simulators via marginal means, with application to neutrino detector design
von: Kim, Hwanwoo, et al.
Veröffentlicht: (2025)
von: Kim, Hwanwoo, et al.
Veröffentlicht: (2025)
Tree of Concepts: Interpretable Continual Learners in Non-Stationary Clinical Domains
von: Cho, Dongkyu, et al.
Veröffentlicht: (2026)
von: Cho, Dongkyu, et al.
Veröffentlicht: (2026)
Average-Reward Maximum Entropy Reinforcement Learning for Underactuated Double Pendulum Tasks
von: Choe, Jean Seong Bjorn, et al.
Veröffentlicht: (2024)
von: Choe, Jean Seong Bjorn, et al.
Veröffentlicht: (2024)
Reward-Zero: Language Embedding Driven Implicit Reward Mechanisms for Reinforcement Learning
von: Zhang, Heng, et al.
Veröffentlicht: (2026)
von: Zhang, Heng, et al.
Veröffentlicht: (2026)
Generalized Gaussian Temporal Difference Error for Uncertainty-aware Reinforcement Learning
von: Kim, Seyeon, et al.
Veröffentlicht: (2024)
von: Kim, Seyeon, et al.
Veröffentlicht: (2024)
Provably Sample-Efficient Robust Reinforcement Learning with Average Reward
von: Roch, Zachary, et al.
Veröffentlicht: (2025)
von: Roch, Zachary, et al.
Veröffentlicht: (2025)
Bandit Simulation for Average Reward Inference
von: Praharaj, Samya, et al.
Veröffentlicht: (2026)
von: Praharaj, Samya, et al.
Veröffentlicht: (2026)
NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models
von: Chong, Hyochan, et al.
Veröffentlicht: (2026)
von: Chong, Hyochan, et al.
Veröffentlicht: (2026)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
Efficient Learning of POMDPs with Known Observation Model in Average-Reward Setting
von: Russo, Alessio, et al.
Veröffentlicht: (2024)
von: Russo, Alessio, et al.
Veröffentlicht: (2024)
Finite-Sample Analysis of Policy Evaluation for Robust Average Reward Reinforcement Learning
von: Xu, Yang, et al.
Veröffentlicht: (2025)
von: Xu, Yang, et al.
Veröffentlicht: (2025)
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
von: Chen, Zijun, et al.
Veröffentlicht: (2025)
von: Chen, Zijun, et al.
Veröffentlicht: (2025)
Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces
von: Kar, Avik, et al.
Veröffentlicht: (2024)
von: Kar, Avik, et al.
Veröffentlicht: (2024)
Neural-network quantum state study of the long-range antiferromagnetic Ising chain
von: Kim, Jicheol, et al.
Veröffentlicht: (2023)
von: Kim, Jicheol, et al.
Veröffentlicht: (2023)
Temporal Difference Learning with Compressed Updates: Error-Feedback meets Reinforcement Learning
von: Mitra, Aritra, et al.
Veröffentlicht: (2023)
von: Mitra, Aritra, et al.
Veröffentlicht: (2023)
Average-Reward Soft Actor-Critic
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
Sample Complexity of Average-Reward Q-Learning: From Single-agent to Federated Reinforcement Learning
von: Jiao, Yuchen, et al.
Veröffentlicht: (2026)
von: Jiao, Yuchen, et al.
Veröffentlicht: (2026)
Optimization of Inter-group Criteria for Clustering with Minimum Size Constraints
von: Laber, Eduardo S., et al.
Veröffentlicht: (2024)
von: Laber, Eduardo S., et al.
Veröffentlicht: (2024)
New bounds on the cohesion of complete-link and other linkage methods for agglomeration clustering
von: Dasgupta, Sanjoy, et al.
Veröffentlicht: (2024)
von: Dasgupta, Sanjoy, et al.
Veröffentlicht: (2024)
Sparse Autoencoders Reveal Temporal Difference Learning in Large Language Models
von: Demircan, Can, et al.
Veröffentlicht: (2024)
von: Demircan, Can, et al.
Veröffentlicht: (2024)
TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Multi-Objective Reinforcement Learning-Guided Diffusion Models for 3D De Novo Molecular Design
von: Chen, Lianghong, et al.
Veröffentlicht: (2025)
von: Chen, Lianghong, et al.
Veröffentlicht: (2025)
Policy Zooming: Adaptive Discretization-based Infinite-Horizon Average-Reward Reinforcement Learning
von: Kar, Avik, et al.
Veröffentlicht: (2024)
von: Kar, Avik, et al.
Veröffentlicht: (2024)
Personalized Multi-Agent Average Reward TD-Learning via Joint Linear Approximation
von: Wang, Leo Muxing, et al.
Veröffentlicht: (2026)
von: Wang, Leo Muxing, et al.
Veröffentlicht: (2026)
A Harmonic Mean Formulation of Average Reward Reinforcement Learning in SMDPs
von: Shtossel, Erel, et al.
Veröffentlicht: (2026)
von: Shtossel, Erel, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Stabilizing Temporal Difference Learning via Implicit Stochastic Recursion
von: Kim, Hwanwoo, et al.
Veröffentlicht: (2025) -
Implicit Q-Learning and SARSA: Liberating Policy Control from Step-Size Calibration
von: Kim, Hwanwoo, et al.
Veröffentlicht: (2026) -
Enhancing Gaussian Process Surrogates for Optimization and Posterior Approximation via Random Exploration
von: Kim, Hwanwoo, et al.
Veröffentlicht: (2024) -
Dealing with the Evil Twins: Improving Random Augmentation by Addressing Catastrophic Forgetting of Diverse Augmentations
von: Cho, Dongkyu, et al.
Veröffentlicht: (2025) -
Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes
von: Blaser, Ethan, et al.
Veröffentlicht: (2026)