Bellman Unbiasedness: Toward Provably Efficient Distributional Reinforcement Learning with General Value Function Approximation
Fuente:
arXiv
Saved in:
| Main Authors: | Cho, Taehyun, Han, Seungyub, Ju, Seokhun, Kim, Dohyeong, Lee, Kyungjae, Lee, Jungwoo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
by: Cho, Taehyun, et al.
Published: (2025)
by: Cho, Taehyun, et al.
Published: (2025)
Spectral-Risk Safe Reinforcement Learning with Convergence Guarantees
by: Kim, Dohyeong, et al.
Published: (2024)
by: Kim, Dohyeong, et al.
Published: (2024)
SPQR: Controlling Q-ensemble Independence with Spiked Random Model for Reinforcement Learning
by: Lee, Dohyeok, et al.
Published: (2024)
by: Lee, Dohyeok, et al.
Published: (2024)
On the Convergence of Continual Learning with Adaptive Methods
by: Han, Seungyub, et al.
Published: (2024)
by: Han, Seungyub, et al.
Published: (2024)
Lyapunov-Guided Self-Alignment: Test-Time Adaptation for Offline Safe Reinforcement Learning
by: Han, Seungyub, et al.
Published: (2026)
by: Han, Seungyub, et al.
Published: (2026)
Learning Generalizable Visuomotor Policy through Dynamics-Alignment
by: Lee, Dohyeok, et al.
Published: (2025)
by: Lee, Dohyeok, et al.
Published: (2025)
Mitigating Resolution-Drift in Federated Learning: Case of Keypoint Detection
by: Lim, Taeheon, et al.
Published: (2025)
by: Lim, Taeheon, et al.
Published: (2025)
Learning Dexterous Grasping from Sparse Taxonomy Guidance
by: Park, Juhan, et al.
Published: (2026)
by: Park, Juhan, et al.
Published: (2026)
The Role of Inherent Bellman Error in Offline Reinforcement Learning with Linear Function Approximation
by: Golowich, Noah, et al.
Published: (2024)
by: Golowich, Noah, et al.
Published: (2024)
Stage-Wise Reward Shaping for Acrobatic Robots: A Constrained Multi-Objective Reinforcement Learning Approach
by: Kim, Dohyeong, et al.
Published: (2024)
by: Kim, Dohyeong, et al.
Published: (2024)
Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection
by: Lee, Kyungjae, et al.
Published: (2024)
by: Lee, Kyungjae, et al.
Published: (2024)
A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
by: Patterson, Andrew, et al.
Published: (2021)
by: Patterson, Andrew, et al.
Published: (2021)
Self-Corrective Task Planning by Inverse Prompting with Large Language Models
by: Lee, Jiho, et al.
Published: (2025)
by: Lee, Jiho, et al.
Published: (2025)
QFlowNet: Fast, Diverse, and Efficient Unitary Synthesis with Generative Flow Networks
by: Koo, Inhoe, et al.
Published: (2026)
by: Koo, Inhoe, et al.
Published: (2026)
Toward Complex-Valued Neural Networks for Waveform Generation
by: Oh, Hyung-Seok, et al.
Published: (2026)
by: Oh, Hyung-Seok, et al.
Published: (2026)
Beyond the Bellman Fixed Point: Geometry and Fast Policy Identification in Value Iteration
by: Lee, Donghwan
Published: (2026)
by: Lee, Donghwan
Published: (2026)
Towards Differentially Private Reinforcement Learning with General Function Approximation
by: He, Yi, et al.
Published: (2026)
by: He, Yi, et al.
Published: (2026)
Efficient Process Reward Modeling via Contrastive Mutual Information
by: Lee, Nakyung, et al.
Published: (2026)
by: Lee, Nakyung, et al.
Published: (2026)
Path-Coupled Bellman Flows for Distributional Reinforcement Learning
by: Xu, Boyang, et al.
Published: (2026)
by: Xu, Boyang, et al.
Published: (2026)
Generalized Gaussian Temporal Difference Error for Uncertainty-aware Reinforcement Learning
by: Kim, Seyeon, et al.
Published: (2024)
by: Kim, Seyeon, et al.
Published: (2024)
Diverse Randomized Value Functions: A Provably Pessimistic Approach for Offline Reinforcement Learning
by: Yu, Xudong, et al.
Published: (2024)
by: Yu, Xudong, et al.
Published: (2024)
KOALA: Empirical Lessons Toward Memory-Efficient and Fast Diffusion Models for Text-to-Image Synthesis
by: Lee, Youngwan, et al.
Published: (2023)
by: Lee, Youngwan, et al.
Published: (2023)
Foundation Model-Driven Framework for Human-Object Interaction Prediction with Segmentation Mask Integration
by: Park, Juhan, et al.
Published: (2025)
by: Park, Juhan, et al.
Published: (2025)
Linear Bellman Completeness Suffices for Efficient Online Reinforcement Learning with Few Actions
by: Golowich, Noah, et al.
Published: (2024)
by: Golowich, Noah, et al.
Published: (2024)
Tensor and Matrix Low-Rank Value-Function Approximation in Reinforcement Learning
by: Rozada, Sergio, et al.
Published: (2022)
by: Rozada, Sergio, et al.
Published: (2022)
Theoretical Barriers in Bellman-Based Reinforcement Learning
by: Pinon, Brieuc, et al.
Published: (2025)
by: Pinon, Brieuc, et al.
Published: (2025)
Provable Distributional Value Iteration under Partial Observability
by: Preuett III, Larry, et al.
Published: (2025)
by: Preuett III, Larry, et al.
Published: (2025)
Sample and Computationally Efficient Continuous-Time Reinforcement Learning with General Function Approximation
by: Zhao, Runze, et al.
Published: (2025)
by: Zhao, Runze, et al.
Published: (2025)
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning
by: Omura, Motoki, et al.
Published: (2025)
by: Omura, Motoki, et al.
Published: (2025)
Analysis of Off-Policy $n$-Step TD-Learning with Linear Function Approximation
by: Lim, Han-Dong, et al.
Published: (2025)
by: Lim, Han-Dong, et al.
Published: (2025)
Provable Reward-Agnostic Preference-Based Reinforcement Learning
by: Zhan, Wenhao, et al.
Published: (2023)
by: Zhan, Wenhao, et al.
Published: (2023)
KU-DMIS at EHRSQL 2024:Generating SQL query via question templatization in EHR
by: Kim, Hajung, et al.
Published: (2024)
by: Kim, Hajung, et al.
Published: (2024)
Understanding the theoretical properties of projected Bellman equation, linear Q-learning, and approximate value iteration
by: Lim, Han-Dong, et al.
Published: (2025)
by: Lim, Han-Dong, et al.
Published: (2025)
Provably Efficient Exploration in Inverse Constrained Reinforcement Learning
by: Yue, Bo, et al.
Published: (2024)
by: Yue, Bo, et al.
Published: (2024)
Adversarial Environment Design via Regret-Guided Diffusion Models
by: Chung, Hojun, et al.
Published: (2024)
by: Chung, Hojun, et al.
Published: (2024)
Provable Offline Reinforcement Learning for Structured Cyclic MDPs
by: Lee, Kyungbok, et al.
Published: (2026)
by: Lee, Kyungbok, et al.
Published: (2026)
Hybrid Deep Searcher: Scalable Parallel and Sequential Search Reasoning
by: Ko, Dayoon, et al.
Published: (2025)
by: Ko, Dayoon, et al.
Published: (2025)
CleaR: Towards Robust and Generalized Parameter-Efficient Fine-Tuning for Noisy Label Learning
by: Kim, Yeachan, et al.
Published: (2024)
by: Kim, Yeachan, et al.
Published: (2024)
When Is Enough Not Enough? Illusory Completion in Search Agents
by: Ko, Dayoon, et al.
Published: (2026)
by: Ko, Dayoon, et al.
Published: (2026)
Mitigating Spurious Correlations via Disagreement Probability
by: Han, Hyeonggeun, et al.
Published: (2024)
by: Han, Hyeonggeun, et al.
Published: (2024)
Similar Items
-
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
by: Cho, Taehyun, et al.
Published: (2025) -
Spectral-Risk Safe Reinforcement Learning with Convergence Guarantees
by: Kim, Dohyeong, et al.
Published: (2024) -
SPQR: Controlling Q-ensemble Independence with Spiked Random Model for Reinforcement Learning
by: Lee, Dohyeok, et al.
Published: (2024) -
On the Convergence of Continual Learning with Adaptive Methods
by: Han, Seungyub, et al.
Published: (2024) -
Lyapunov-Guided Self-Alignment: Test-Time Adaptation for Offline Safe Reinforcement Learning
by: Han, Seungyub, et al.
Published: (2026)