An Analysis of Action-Value Temporal-Difference Methods That Learn State Values
Fuente:
arXiv
Saved in:
| Main Authors: | Daley, Brett, Nagarajan, Prabhat, White, Martha, Machado, Marlos C. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Demystifying the Recency Heuristic in Temporal-Difference Learning
by: Daley, Brett, et al.
Published: (2024)
by: Daley, Brett, et al.
Published: (2024)
Deep Double Q-learning
by: Nagarajan, Prabhat, et al.
Published: (2025)
by: Nagarajan, Prabhat, et al.
Published: (2025)
Deep Reinforcement Learning with Gradient Eligibility Traces
by: Elelimy, Esraa, et al.
Published: (2025)
by: Elelimy, Esraa, et al.
Published: (2025)
Averaging $n$-step Returns Reduces Variance in Reinforcement Learning
by: Daley, Brett, et al.
Published: (2024)
by: Daley, Brett, et al.
Published: (2024)
When is Offline Policy Selection Sample Efficient for Reinforcement Learning?
by: Liu, Vincent, et al.
Published: (2023)
by: Liu, Vincent, et al.
Published: (2023)
Trajectory-Aware Eligibility Traces for Off-Policy Reinforcement Learning
by: Daley, Brett, et al.
Published: (2023)
by: Daley, Brett, et al.
Published: (2023)
Harnessing Discrete Representations For Continual Reinforcement Learning
by: Meyer, Edan, et al.
Published: (2023)
by: Meyer, Edan, et al.
Published: (2023)
AGaLiTe: Approximate Gated Linear Transformers for Online Reinforcement Learning
by: Pramanik, Subhojeet, et al.
Published: (2023)
by: Pramanik, Subhojeet, et al.
Published: (2023)
A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
by: Patterson, Andrew, et al.
Published: (2021)
by: Patterson, Andrew, et al.
Published: (2021)
Value Bonuses using Ensemble Errors for Exploration in Reinforcement Learning
by: Wahab, Abdul, et al.
Published: (2026)
by: Wahab, Abdul, et al.
Published: (2026)
Gradient Iterated Temporal-Difference Learning
by: Vincent, Théo, et al.
Published: (2026)
by: Vincent, Théo, et al.
Published: (2026)
Proper Laplacian Representation Learning
by: Gomez, Diego, et al.
Published: (2023)
by: Gomez, Diego, et al.
Published: (2023)
The Laplacian Keyboard: Beyond the Linear Span
by: Chandrasekar, Siddarth, et al.
Published: (2026)
by: Chandrasekar, Siddarth, et al.
Published: (2026)
A Study of Value-Aware Eigenoptions
by: Kotamreddy, Harshil, et al.
Published: (2025)
by: Kotamreddy, Harshil, et al.
Published: (2025)
Mitigating Value Hallucination in Dyna Planning via Multistep Predecessor Models
by: Aminmansour, Farzane, et al.
Published: (2020)
by: Aminmansour, Farzane, et al.
Published: (2020)
Distributions as Actions: A Unified Framework for Diverse Action Spaces
by: He, Jiamin, et al.
Published: (2025)
by: He, Jiamin, et al.
Published: (2025)
Guided Flow Policy: Learning from High-Value Actions in Offline Reinforcement Learning
by: Tiofack, Franki Nguimatsia, et al.
Published: (2025)
by: Tiofack, Franki Nguimatsia, et al.
Published: (2025)
Beyond Shapley Values: Cooperative Games for the Interpretation of Machine Learning Models
by: Idrissi, Marouane Il, et al.
Published: (2025)
by: Idrissi, Marouane Il, et al.
Published: (2025)
Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning
by: Ahn, Hongjoon, et al.
Published: (2025)
by: Ahn, Hongjoon, et al.
Published: (2025)
Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration
by: Kim, Dongyoung, et al.
Published: (2023)
by: Kim, Dongyoung, et al.
Published: (2023)
Is there Value in Reinforcement Learning?
by: Fox, Lior, et al.
Published: (2025)
by: Fox, Lior, et al.
Published: (2025)
A Look at Value-Based Decision-Time vs. Background Planning Methods Across Different Settings
by: Alver, Safa, et al.
Published: (2022)
by: Alver, Safa, et al.
Published: (2022)
GeoMAE: Masking Representation Learning for Spatio-Temporal Graph Forecasting with Missing Values
by: Ke, Songyu, et al.
Published: (2025)
by: Ke, Songyu, et al.
Published: (2025)
The Cell Must Go On: Agar.io for Continual Reinforcement Learning
by: Mohamed, Mohamed A., et al.
Published: (2025)
by: Mohamed, Mohamed A., et al.
Published: (2025)
Temporal-Difference Variational Continual Learning
by: Melo, Luckeciano C., et al.
Published: (2024)
by: Melo, Luckeciano C., et al.
Published: (2024)
Learning from Disagreement: Clinician Overrides as Implicit Preference Signals for Clinical AI in Value-Based Care
by: Singh, Prabhjot, et al.
Published: (2026)
by: Singh, Prabhjot, et al.
Published: (2026)
Adaptive Action Chunking via Multi-Chunk Q Value Estimation
by: Shin, Yongjae, et al.
Published: (2026)
by: Shin, Yongjae, et al.
Published: (2026)
Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration
by: Chen, Yan, et al.
Published: (2025)
by: Chen, Yan, et al.
Published: (2025)
Finite-Time Analysis of Temporal Difference Learning with Experience Replay
by: Lim, Han-Dong, et al.
Published: (2023)
by: Lim, Han-Dong, et al.
Published: (2023)
Real-Time Recurrent Learning using Trace Units in Reinforcement Learning
by: Elelimy, Esraa, et al.
Published: (2024)
by: Elelimy, Esraa, et al.
Published: (2024)
Expected Possession Value of Control and Duel Actions for Soccer Player's Skills Estimation
by: Shelopugin, Andrei
Published: (2024)
by: Shelopugin, Andrei
Published: (2024)
Active Inference with Reusable State-Dependent Value Profiles
by: Poschl, Jacob
Published: (2025)
by: Poschl, Jacob
Published: (2025)
POWQMIX: Weighted Value Factorization with Potentially Optimal Joint Actions Recognition for Cooperative Multi-Agent Reinforcement Learning
by: Huang, Chang, et al.
Published: (2024)
by: Huang, Chang, et al.
Published: (2024)
Empirical Design in Reinforcement Learning
by: Patterson, Andrew, et al.
Published: (2023)
by: Patterson, Andrew, et al.
Published: (2023)
Discerning Temporal Difference Learning
by: Ma, Jianfei
Published: (2023)
by: Ma, Jianfei
Published: (2023)
Backstepping Temporal Difference Learning
by: Lim, Han-Dong, et al.
Published: (2023)
by: Lim, Han-Dong, et al.
Published: (2023)
PAWN: Piece Value Analysis with Neural Networks
by: Tang, Ethan, et al.
Published: (2026)
by: Tang, Ethan, et al.
Published: (2026)
Finite Sample Analysis of Linear Temporal Difference Learning with Arbitrary Features
by: Xie, Zixuan, et al.
Published: (2025)
by: Xie, Zixuan, et al.
Published: (2025)
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
by: Huang, Chenghua, et al.
Published: (2025)
by: Huang, Chenghua, et al.
Published: (2025)
VDSC: Enhancing Exploration Timing with Value Discrepancy and State Counts
by: Captari, Marius, et al.
Published: (2024)
by: Captari, Marius, et al.
Published: (2024)
Similar Items
-
Demystifying the Recency Heuristic in Temporal-Difference Learning
by: Daley, Brett, et al.
Published: (2024) -
Deep Double Q-learning
by: Nagarajan, Prabhat, et al.
Published: (2025) -
Deep Reinforcement Learning with Gradient Eligibility Traces
by: Elelimy, Esraa, et al.
Published: (2025) -
Averaging $n$-step Returns Reduces Variance in Reinforcement Learning
by: Daley, Brett, et al.
Published: (2024) -
When is Offline Policy Selection Sample Efficient for Reinforcement Learning?
by: Liu, Vincent, et al.
Published: (2023)