Online Estimation and Inference for Robust Policy Evaluation in Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Weidong, Tu, Jiyuan, Chen, Xi, Zhang, Yichen |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distributed Estimation and Inference for Semi-parametric Binary Response Models
by: Chen, Xi, et al.
Published: (2022)
by: Chen, Xi, et al.
Published: (2022)
Robust Variational Bayes by Min-Max Median Aggregation
by: Yan, Jiawei, et al.
Published: (2025)
by: Yan, Jiawei, et al.
Published: (2025)
Majority Vote for Distributed Differentially Private Sign Selection
by: Liu, Weidong, et al.
Published: (2022)
by: Liu, Weidong, et al.
Published: (2022)
Improving Policy Exploitation in Online Reinforcement Learning with Instant Retrospect Action
by: Gao, Gong, et al.
Published: (2026)
by: Gao, Gong, et al.
Published: (2026)
Acceleration of stochastic gradient descent with momentum by averaging: finite-sample rates and asymptotic normality
by: Tang, Kejie, et al.
Published: (2023)
by: Tang, Kejie, et al.
Published: (2023)
Online Statistical Inference for Contextual Bandits via Stochastic Gradient Descent
by: Chang, Xiangyu, et al.
Published: (2022)
by: Chang, Xiangyu, et al.
Published: (2022)
Online Tensor Inference
by: Wen, Xin, et al.
Published: (2023)
by: Wen, Xin, et al.
Published: (2023)
Doubly Robust Interval Estimation for Optimal Policy Evaluation in Online Learning
by: Shen, Ye, et al.
Published: (2021)
by: Shen, Ye, et al.
Published: (2021)
Doubly Optimal Policy Evaluation for Reinforcement Learning
by: Liu, Shuze Daniel, et al.
Published: (2024)
by: Liu, Shuze Daniel, et al.
Published: (2024)
Efficient Multi-Policy Evaluation for Reinforcement Learning
by: Liu, Shuze Daniel, et al.
Published: (2024)
by: Liu, Shuze Daniel, et al.
Published: (2024)
Adaptive Estimation and Inference in Conditional Moment Models via the Discrepancy Principle
by: Tan, Jiyuan, et al.
Published: (2026)
by: Tan, Jiyuan, et al.
Published: (2026)
Efficient Policy Evaluation with Safety Constraint for Reinforcement Learning
by: Chen, Claire, et al.
Published: (2024)
by: Chen, Claire, et al.
Published: (2024)
Estimation and Inference in Distributional Reinforcement Learning
by: Zhang, Liangyu, et al.
Published: (2023)
by: Zhang, Liangyu, et al.
Published: (2023)
Toward Evaluating Robustness of Reinforcement Learning with Adversarial Policy
by: Zheng, Xiang, et al.
Published: (2023)
by: Zheng, Xiang, et al.
Published: (2023)
Efficient Online Reinforcement Learning for Diffusion Policy
by: Ma, Haitong, et al.
Published: (2025)
by: Ma, Haitong, et al.
Published: (2025)
Online Policy Learning and Inference by Matrix Completion
by: Duan, Congyuan, et al.
Published: (2024)
by: Duan, Congyuan, et al.
Published: (2024)
Online Statistical Inference in Decision-Making with Matrix Context
by: Han, Qiyu, et al.
Published: (2022)
by: Han, Qiyu, et al.
Published: (2022)
Evolving Diffusion and Flow Matching Policies for Online Reinforcement Learning
by: Zhang, Chubin, et al.
Published: (2025)
by: Zhang, Chubin, et al.
Published: (2025)
Policy-Aware Design of Large-Scale Factorial Experiments
by: Wen, Xin, et al.
Published: (2026)
by: Wen, Xin, et al.
Published: (2026)
Doubly-Robust Off-Policy Evaluation with Estimated Logging Policy
by: Lee, Kyungbok, et al.
Published: (2024)
by: Lee, Kyungbok, et al.
Published: (2024)
Blending Imitation and Reinforcement Learning for Robust Policy Improvement
by: Liu, Xuefeng, et al.
Published: (2023)
by: Liu, Xuefeng, et al.
Published: (2023)
Finite-Sample Analysis of Policy Evaluation for Robust Average Reward Reinforcement Learning
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Iterative Refinement of Flow Policies in Probability Space for Online Reinforcement Learning
by: Sun, Mingyang, et al.
Published: (2025)
by: Sun, Mingyang, et al.
Published: (2025)
Maximum Entropy Reinforcement Learning with Diffusion Policy
by: Dong, Xiaoyi, et al.
Published: (2025)
by: Dong, Xiaoyi, et al.
Published: (2025)
Flow-Based Policy for Online Reinforcement Learning
by: Lv, Lei, et al.
Published: (2025)
by: Lv, Lei, et al.
Published: (2025)
Online Robust Reinforcement Learning with General Function Approximation
by: Ghosh, Debamita, et al.
Published: (2025)
by: Ghosh, Debamita, et al.
Published: (2025)
Learning by Doing: An Online Causal Reinforcement Learning Framework with Causal-Aware Policy
by: Cai, Ruichu, et al.
Published: (2024)
by: Cai, Ruichu, et al.
Published: (2024)
Towards Robust Offline-to-Online Reinforcement Learning via Uncertainty and Smoothness
by: Wen, Xiaoyu, et al.
Published: (2023)
by: Wen, Xiaoyu, et al.
Published: (2023)
SGD with Dependent Data: Optimal Estimation, Regret, and Inference
by: Shen, Yinan, et al.
Published: (2026)
by: Shen, Yinan, et al.
Published: (2026)
Towards Fast Safe Online Reinforcement Learning via Policy Finetuning
by: Chen, Keru, et al.
Published: (2024)
by: Chen, Keru, et al.
Published: (2024)
Energy-Guided Diffusion Sampling for Offline-to-Online Reinforcement Learning
by: Liu, Xu-Hui, et al.
Published: (2024)
by: Liu, Xu-Hui, et al.
Published: (2024)
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning
by: Zhang, Tonghe, et al.
Published: (2025)
by: Zhang, Tonghe, et al.
Published: (2025)
ORVIT: Near-Optimal Online Distributionally Robust Reinforcement Learning
by: Ghosh, Debamita, et al.
Published: (2025)
by: Ghosh, Debamita, et al.
Published: (2025)
Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference
by: Zhang, Qining, et al.
Published: (2024)
by: Zhang, Qining, et al.
Published: (2024)
Online Matching via Reinforcement Learning: An Expert Policy Orchestration Strategy
by: Mignacco, Chiara, et al.
Published: (2025)
by: Mignacco, Chiara, et al.
Published: (2025)
A Non-Monolithic Policy Approach of Offline-to-Online Reinforcement Learning
by: Kim, JaeYoon, et al.
Published: (2024)
by: Kim, JaeYoon, et al.
Published: (2024)
RLOMM: An Efficient and Robust Online Map Matching Framework with Reinforcement Learning
by: Chen, Minxiao, et al.
Published: (2025)
by: Chen, Minxiao, et al.
Published: (2025)
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
by: Zhang, Zijing, et al.
Published: (2025)
by: Zhang, Zijing, et al.
Published: (2025)
HCPO: Hierarchical Conductor-Based Policy Optimization in Multi-Agent Reinforcement Learning
by: Liu, Zejiao, et al.
Published: (2025)
by: Liu, Zejiao, et al.
Published: (2025)
Settling the Sample Complexity of Online Reinforcement Learning
by: Zhang, Zihan, et al.
Published: (2023)
by: Zhang, Zihan, et al.
Published: (2023)
Similar Items
-
Distributed Estimation and Inference for Semi-parametric Binary Response Models
by: Chen, Xi, et al.
Published: (2022) -
Robust Variational Bayes by Min-Max Median Aggregation
by: Yan, Jiawei, et al.
Published: (2025) -
Majority Vote for Distributed Differentially Private Sign Selection
by: Liu, Weidong, et al.
Published: (2022) -
Improving Policy Exploitation in Online Reinforcement Learning with Instant Retrospect Action
by: Gao, Gong, et al.
Published: (2026) -
Acceleration of stochastic gradient descent with momentum by averaging: finite-sample rates and asymptotic normality
by: Tang, Kejie, et al.
Published: (2023)