Estimation and Inference in Distributional Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Liangyu, Peng, Yang, Liang, Jiadong, Yang, Wenhao, Zhang, Zhihua |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Statistical Efficiency of Distributional Temporal Difference Learning and Freedman's Inequality in Hilbert Spaces
by: Peng, Yang, et al.
Published: (2024)
by: Peng, Yang, et al.
Published: (2024)
A Finite Sample Analysis of Distributional TD Learning with Linear Function Approximation
by: Peng, Yang, et al.
Published: (2025)
by: Peng, Yang, et al.
Published: (2025)
Federated Reinforcement Learning with Constraint Heterogeneity
by: Jin, Hao, et al.
Published: (2024)
by: Jin, Hao, et al.
Published: (2024)
Federated Control in Markov Decision Processes
by: Jin, Hao, et al.
Published: (2024)
by: Jin, Hao, et al.
Published: (2024)
Asymptotic Time-Uniform Inference for Parameters in Averaged Stochastic Approximation
by: Xie, Chuhan, et al.
Published: (2024)
by: Xie, Chuhan, et al.
Published: (2024)
Accelerated Distributional Temporal Difference Learning with Linear Function Approximation
by: Jin, Kaicheng, et al.
Published: (2025)
by: Jin, Kaicheng, et al.
Published: (2025)
Decoupled Functional Central Limit Theorems for Two-Time-Scale Stochastic Approximation
by: Han, Yuze, et al.
Published: (2024)
by: Han, Yuze, et al.
Published: (2024)
OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning
by: Yao, Yihang, et al.
Published: (2024)
by: Yao, Yihang, et al.
Published: (2024)
Online Estimation and Inference for Robust Policy Evaluation in Reinforcement Learning
by: Liu, Weidong, et al.
Published: (2023)
by: Liu, Weidong, et al.
Published: (2023)
Towards Monotonic Improvement in In-Context Reinforcement Learning
by: Zhang, Wenhao, et al.
Published: (2025)
by: Zhang, Wenhao, et al.
Published: (2025)
Flow-based Policy With Distributional Reinforcement Learning in Trajectory Optimization
by: Hao, Ruijie, et al.
Published: (2026)
by: Hao, Ruijie, et al.
Published: (2026)
Distributed Online Convex Optimization with Efficient Communication: Improved Algorithm and Lower bounds
by: Yang, Sifan, et al.
Published: (2026)
by: Yang, Sifan, et al.
Published: (2026)
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
by: Panaganti, Kishan, et al.
Published: (2026)
by: Panaganti, Kishan, et al.
Published: (2026)
Quantile Geometry Regularization for Distributional Reinforcement Learning
by: Zhang, Zhaofan, et al.
Published: (2026)
by: Zhang, Zhaofan, et al.
Published: (2026)
Single-Trajectory Distributionally Robust Reinforcement Learning
by: Liang, Zhipeng, et al.
Published: (2023)
by: Liang, Zhipeng, et al.
Published: (2023)
Improve the Training Efficiency of DRL for Wireless Communication Resource Allocation: The Role of Generative Diffusion Models
by: Zhang, Xinren, et al.
Published: (2025)
by: Zhang, Xinren, et al.
Published: (2025)
PAHQ: Accelerating Automated Circuit Discovery through Mixed-Precision Inference Optimization
by: Wang, Xinhai, et al.
Published: (2025)
by: Wang, Xinhai, et al.
Published: (2025)
Optimal Multi-Distribution Learning
by: Zhang, Zihan, et al.
Published: (2023)
by: Zhang, Zihan, et al.
Published: (2023)
Discounted Online Convex Optimization: Uniform Regret Across a Continuous Interval
by: Yang, Wenhao, et al.
Published: (2025)
by: Yang, Wenhao, et al.
Published: (2025)
Machine Learning-Assisted High-Dimensional Matrix Estimation
by: Tian, Wan, et al.
Published: (2026)
by: Tian, Wan, et al.
Published: (2026)
Robust Bandwidth Estimation for Real-Time Communication with Offline Reinforcement Learning
by: Kai, Jian, et al.
Published: (2025)
by: Kai, Jian, et al.
Published: (2025)
Distributed Estimation and Inference for Semi-parametric Binary Response Models
by: Chen, Xi, et al.
Published: (2022)
by: Chen, Xi, et al.
Published: (2022)
Fast Stochastic Policy Gradient: Negative Momentum for Reinforcement Learning
by: Zhang, Haobin, et al.
Published: (2024)
by: Zhang, Haobin, et al.
Published: (2024)
Statistical Inference of the Value Function for Reinforcement Learning in Infinite Horizon Settings
by: Shi, C., et al.
Published: (2020)
by: Shi, C., et al.
Published: (2020)
Ranking-Aware Calibration for Reliable Multimodal Reinforcement Learning
by: Cui, Peng, et al.
Published: (2026)
by: Cui, Peng, et al.
Published: (2026)
Universal Online Convex Optimization with $1$ Projection per Round
by: Yang, Wenhao, et al.
Published: (2024)
by: Yang, Wenhao, et al.
Published: (2024)
Offline Multi-agent Reinforcement Learning via Sequential Score Decomposition
by: Qiao, Dan, et al.
Published: (2025)
by: Qiao, Dan, et al.
Published: (2025)
Robust Offline Reinforcement Learning for Non-Markovian Decision Processes
by: Huang, Ruiquan, et al.
Published: (2024)
by: Huang, Ruiquan, et al.
Published: (2024)
Low-dimensional adaptation of diffusion models: Convergence in total variation
by: Liang, Jiadong, et al.
Published: (2025)
by: Liang, Jiadong, et al.
Published: (2025)
RL-MUL 2.0: Multiplier Design Optimization with Parallel Deep Reinforcement Learning and Space Reduction
by: Zuo, Dongsheng, et al.
Published: (2024)
by: Zuo, Dongsheng, et al.
Published: (2024)
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM
by: Duong, Thang, et al.
Published: (2025)
by: Duong, Thang, et al.
Published: (2025)
Enhance the Safety in Reinforcement Learning by ADRC Lagrangian Methods
by: Zhang, Mingxu, et al.
Published: (2026)
by: Zhang, Mingxu, et al.
Published: (2026)
Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference
by: Zhang, Qining, et al.
Published: (2024)
by: Zhang, Qining, et al.
Published: (2024)
Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling
by: Qi, Han, et al.
Published: (2025)
by: Qi, Han, et al.
Published: (2025)
Nearly Optimal Bayesian Inference for Structural Missingness
by: Liang, Chen, et al.
Published: (2026)
by: Liang, Chen, et al.
Published: (2026)
Privacy-Aware Multi-Device Cooperative Edge Inference with Distributed Resource Bidding
by: Zhuang, Wenhao, et al.
Published: (2024)
by: Zhuang, Wenhao, et al.
Published: (2024)
Reinforcement Learning for Machine Learning Engineering Agents
by: Yang, Sherry, et al.
Published: (2025)
by: Yang, Sherry, et al.
Published: (2025)
Post Reinforcement Learning Inference
by: Syrgkanis, Vasilis, et al.
Published: (2023)
by: Syrgkanis, Vasilis, et al.
Published: (2023)
DistPred: A Distribution-Free Probabilistic Inference Method for Regression and Forecasting
by: Liang, Daojun, et al.
Published: (2024)
by: Liang, Daojun, et al.
Published: (2024)
Accelerating Sparse Transformer Inference on GPU
by: Dai, Wenhao, et al.
Published: (2025)
by: Dai, Wenhao, et al.
Published: (2025)
Similar Items
-
Statistical Efficiency of Distributional Temporal Difference Learning and Freedman's Inequality in Hilbert Spaces
by: Peng, Yang, et al.
Published: (2024) -
A Finite Sample Analysis of Distributional TD Learning with Linear Function Approximation
by: Peng, Yang, et al.
Published: (2025) -
Federated Reinforcement Learning with Constraint Heterogeneity
by: Jin, Hao, et al.
Published: (2024) -
Federated Control in Markov Decision Processes
by: Jin, Hao, et al.
Published: (2024) -
Asymptotic Time-Uniform Inference for Parameters in Averaged Stochastic Approximation
by: Xie, Chuhan, et al.
Published: (2024)