Low Variance Off-policy Evaluation with State-based Importance Sampling
Fuente:
arXiv
Saved in:
| Main Authors: | Bossens, David M., Thomas, Philip S. |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust Lagrangian and Adversarial Policy Gradient for Robust Constrained Markov Decision Processes
by: Bossens, David M.
Published: (2023)
by: Bossens, David M.
Published: (2023)
Quotient DAGs for Off-Policy Evaluation:Forward-Flow Importance Sampling and Exact Slate Propensities
by: Xie, Ziwen, et al.
Published: (2026)
by: Xie, Ziwen, et al.
Published: (2026)
Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation
by: Zhou, Hongyi, et al.
Published: (2025)
by: Zhou, Hongyi, et al.
Published: (2025)
Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation
by: Chaudhari, Shreyas, et al.
Published: (2024)
by: Chaudhari, Shreyas, et al.
Published: (2024)
Model Selection for Off-policy Evaluation: New Algorithms and Experimental Protocol
by: Liu, Pai, et al.
Published: (2025)
by: Liu, Pai, et al.
Published: (2025)
Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs
by: Huang, Luke J., et al.
Published: (2026)
by: Huang, Luke J., et al.
Published: (2026)
Statistical Tractability of Off-policy Evaluation of History-dependent Policies in POMDPs
by: Zhang, Yuheng, et al.
Published: (2025)
by: Zhang, Yuheng, et al.
Published: (2025)
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration
by: Zhao, Heyang, et al.
Published: (2025)
by: Zhao, Heyang, et al.
Published: (2025)
Causal Deepsets for Off-policy Evaluation under Spatial or Spatio-temporal Interferences
by: Dai, Runpeng, et al.
Published: (2024)
by: Dai, Runpeng, et al.
Published: (2024)
On the Curses of Future and History in Future-dependent Value Functions for Off-policy Evaluation
by: Zhang, Yuheng, et al.
Published: (2024)
by: Zhang, Yuheng, et al.
Published: (2024)
Importance Sampling for Nonlinear Models
by: Rajmohan, Prakash Palanivelu, et al.
Published: (2025)
by: Rajmohan, Prakash Palanivelu, et al.
Published: (2025)
Primal-Dual Spectral Representation for Off-policy Evaluation
by: Hu, Yang, et al.
Published: (2024)
by: Hu, Yang, et al.
Published: (2024)
Learning Off-policy with Model-based Intrinsic Motivation For Active Online Exploration
by: Wang, Yibo, et al.
Published: (2024)
by: Wang, Yibo, et al.
Published: (2024)
Partial Domain Adaptation via Importance Sampling-based Shift Correction
by: Guo, Cheng-Jun, et al.
Published: (2025)
by: Guo, Cheng-Jun, et al.
Published: (2025)
Quantifying Variance in Evaluation Benchmarks
by: Madaan, Lovish, et al.
Published: (2024)
by: Madaan, Lovish, et al.
Published: (2024)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
by: Lee, Haanvid, et al.
Published: (2024)
by: Lee, Haanvid, et al.
Published: (2024)
A Unifying View of Coverage in Linear Off-Policy Evaluation
by: Amortila, Philip, et al.
Published: (2026)
by: Amortila, Philip, et al.
Published: (2026)
Measuring Sample Importance in Data Pruning for Language Models based on Information Entropy
by: Kim, Minsang, et al.
Published: (2024)
by: Kim, Minsang, et al.
Published: (2024)
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
by: Goodall, Alexander W., et al.
Published: (2025)
by: Goodall, Alexander W., et al.
Published: (2025)
Bootstrap Off-policy with World Model
by: Zhan, Guojian, et al.
Published: (2025)
by: Zhan, Guojian, et al.
Published: (2025)
ESPO: Entropy Importance Sampling Policy Optimization
by: Sheng, Yuepeng, et al.
Published: (2025)
by: Sheng, Yuepeng, et al.
Published: (2025)
AIS: Adaptive Importance Sampling for Quantized RL
by: Zhou, Jiajun, et al.
Published: (2026)
by: Zhou, Jiajun, et al.
Published: (2026)
GIPO: Gaussian Importance Sampling Policy Optimization
by: Lu, Chengxuan, et al.
Published: (2026)
by: Lu, Chengxuan, et al.
Published: (2026)
Variance-Reduction Guidance: Sampling Trajectory Optimization for Diffusion Models
by: Xu, Shifeng, et al.
Published: (2025)
by: Xu, Shifeng, et al.
Published: (2025)
Failure Probability Estimation for Black-Box Autonomous Systems using State-Dependent Importance Sampling Proposals
by: Delecki, Harrison, et al.
Published: (2024)
by: Delecki, Harrison, et al.
Published: (2024)
Imitation from Diverse Behaviors: Wasserstein Quality Diversity Imitation Learning with Single-Step Archive Exploration
by: Yu, Xingrui, et al.
Published: (2024)
by: Yu, Xingrui, et al.
Published: (2024)
Preference as Reward, Maximum Preference Optimization with Importance Sampling
by: Jiang, Zaifan, et al.
Published: (2023)
by: Jiang, Zaifan, et al.
Published: (2023)
Beyond Importance Sampling: Rejection-Gated Policy Optimization
by: Sun, Ziwu, et al.
Published: (2026)
by: Sun, Ziwu, et al.
Published: (2026)
Unified PAC-Bayesian Study of Pessimism for Offline Policy Learning with Regularized Importance Sampling
by: Aouali, Imad, et al.
Published: (2024)
by: Aouali, Imad, et al.
Published: (2024)
Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
by: Li, Chenliang, et al.
Published: (2025)
by: Li, Chenliang, et al.
Published: (2025)
VADE: Variance-Aware Dynamic Sampling via Online Sample-Level Difficulty Estimation for Multimodal RL
by: Hu, Zengjie, et al.
Published: (2025)
by: Hu, Zengjie, et al.
Published: (2025)
Bayesian Off-Policy Evaluation and Learning for Large Action Spaces
by: Aouali, Imad, et al.
Published: (2024)
by: Aouali, Imad, et al.
Published: (2024)
Clustering Context in Off-Policy Evaluation
by: Guzman-Olivares, Daniel, et al.
Published: (2025)
by: Guzman-Olivares, Daniel, et al.
Published: (2025)
Concept-driven Off Policy Evaluation
by: Majumdar, Ritam, et al.
Published: (2024)
by: Majumdar, Ritam, et al.
Published: (2024)
ExO-PPO: an Extended Off-policy Proximal Policy Optimization Algorithm
by: Wang, Hanyong, et al.
Published: (2026)
by: Wang, Hanyong, et al.
Published: (2026)
Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function Approximation
by: Che, Fengdi, et al.
Published: (2024)
by: Che, Fengdi, et al.
Published: (2024)
MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
by: Liu, Yuxi, et al.
Published: (2025)
by: Liu, Yuxi, et al.
Published: (2025)
A Framework for Fair Evaluation of Variance-Aware Bandit Algorithms
by: Wolf, Elise
Published: (2025)
by: Wolf, Elise
Published: (2025)
Uncertainty Estimation using Variance-Gated Distributions
by: Gillis, H. Martin, et al.
Published: (2025)
by: Gillis, H. Martin, et al.
Published: (2025)
A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
by: Patterson, Andrew, et al.
Published: (2021)
by: Patterson, Andrew, et al.
Published: (2021)
Similar Items
-
Robust Lagrangian and Adversarial Policy Gradient for Robust Constrained Markov Decision Processes
by: Bossens, David M.
Published: (2023) -
Quotient DAGs for Off-Policy Evaluation:Forward-Flow Importance Sampling and Exact Slate Propensities
by: Xie, Ziwen, et al.
Published: (2026) -
Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation
by: Zhou, Hongyi, et al.
Published: (2025) -
Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation
by: Chaudhari, Shreyas, et al.
Published: (2024) -
Model Selection for Off-policy Evaluation: New Algorithms and Experimental Protocol
by: Liu, Pai, et al.
Published: (2025)