Efficient Policy Evaluation with Offline Data Informed Behavior Policy Design
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Shuze, Zhang, Shangtong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Efficient Multi-Policy Evaluation for Reinforcement Learning
por: Liu, Shuze Daniel, et al.
Publicado: (2024)
por: Liu, Shuze Daniel, et al.
Publicado: (2024)
Efficient Policy Evaluation with Safety Constraint for Reinforcement Learning
por: Chen, Claire, et al.
Publicado: (2024)
por: Chen, Claire, et al.
Publicado: (2024)
Doubly Optimal Policy Evaluation for Reinforcement Learning
por: Liu, Shuze Daniel, et al.
Publicado: (2024)
por: Liu, Shuze Daniel, et al.
Publicado: (2024)
The ODE Method for Stochastic Approximation and Reinforcement Learning with Markovian Noise
por: Liu, Shuze Daniel, et al.
Publicado: (2024)
por: Liu, Shuze Daniel, et al.
Publicado: (2024)
Convergence of Two-Timescale Markovian Stochastic Approximations with Applications in Reinforcement Learning
por: Mahadevan, Vagul, et al.
Publicado: (2026)
por: Mahadevan, Vagul, et al.
Publicado: (2026)
Adaptive Policy Selection and Fine-Tuning under Interaction Budgets for Offline-to-Online Reinforcement Learning
por: Bozkurt, Alper Kamil, et al.
Publicado: (2026)
por: Bozkurt, Alper Kamil, et al.
Publicado: (2026)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
por: Xie, Zixuan, et al.
Publicado: (2026)
por: Xie, Zixuan, et al.
Publicado: (2026)
Predicting Plasticity in Deep Continual Learning: A Theoretical Perspective
por: Wang, Jiuqi, et al.
Publicado: (2026)
por: Wang, Jiuqi, et al.
Publicado: (2026)
Global Optimality and Finite Sample Analysis of Softmax Off-Policy Actor Critic under State Distribution Mismatch
por: Zhang, Shangtong, et al.
Publicado: (2021)
por: Zhang, Shangtong, et al.
Publicado: (2021)
Constrained Policy Optimization with Explicit Behavior Density for Offline Reinforcement Learning
por: Zhang, Jing, et al.
Publicado: (2023)
por: Zhang, Jing, et al.
Publicado: (2023)
Offline Policy Evaluation for Reinforcement Learning with Adaptively Collected Data
por: Madhow, Sunil, et al.
Publicado: (2023)
por: Madhow, Sunil, et al.
Publicado: (2023)
Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning
por: Gao, Chen-Xiao, et al.
Publicado: (2025)
por: Gao, Chen-Xiao, et al.
Publicado: (2025)
EXTRACT: Efficient Policy Learning by Extracting Transferable Robot Skills from Offline Data
por: Zhang, Jesse, et al.
Publicado: (2024)
por: Zhang, Jesse, et al.
Publicado: (2024)
Revisiting a Design Choice in Gradient Temporal Difference Learning
por: Qian, Xiaochi, et al.
Publicado: (2023)
por: Qian, Xiaochi, et al.
Publicado: (2023)
Mildly Constrained Evaluation Policy for Offline Reinforcement Learning
por: Xu, Linjie, et al.
Publicado: (2023)
por: Xu, Linjie, et al.
Publicado: (2023)
Distributional Offline Policy Evaluation with Predictive Error Guarantees
por: Wu, Runzhe, et al.
Publicado: (2023)
por: Wu, Runzhe, et al.
Publicado: (2023)
Distorted Distributional Policy Evaluation for Offline Reinforcement Learning
por: Iwaki, Ryo, et al.
Publicado: (2026)
por: Iwaki, Ryo, et al.
Publicado: (2026)
Diffusion Policies for Risk-Averse Behavior Modeling in Offline Reinforcement Learning
por: Chen, Xiaocong, et al.
Publicado: (2024)
por: Chen, Xiaocong, et al.
Publicado: (2024)
When is Offline Policy Selection Sample Efficient for Reinforcement Learning?
por: Liu, Vincent, et al.
Publicado: (2023)
por: Liu, Vincent, et al.
Publicado: (2023)
Towards Formalizing Reinforcement Learning Theory
por: Zhang, Shangtong
Publicado: (2025)
por: Zhang, Shangtong
Publicado: (2025)
Online Policy Learning from Offline Preferences
por: Zhang, Guoxi, et al.
Publicado: (2024)
por: Zhang, Guoxi, et al.
Publicado: (2024)
Learning Control Policies for Variable Objectives from Offline Data
por: Weber, Marc, et al.
Publicado: (2023)
por: Weber, Marc, et al.
Publicado: (2023)
Fat-to-Thin Policy Optimization: Offline RL with Sparse Policies
por: Zhu, Lingwei, et al.
Publicado: (2025)
por: Zhu, Lingwei, et al.
Publicado: (2025)
Evaluation-Time Policy Switching for Offline Reinforcement Learning
por: Neggatu, Natinael Solomon, et al.
Publicado: (2025)
por: Neggatu, Natinael Solomon, et al.
Publicado: (2025)
Sample-Efficient Policy Constraint Offline Deep Reinforcement Learning based on Sample Filtering
por: Chen, Yuanhao, et al.
Publicado: (2025)
por: Chen, Yuanhao, et al.
Publicado: (2025)
Robust Offline Policy Learning with Observational Data from Multiple Sources
por: Carranza, Aldo Gael, et al.
Publicado: (2024)
por: Carranza, Aldo Gael, et al.
Publicado: (2024)
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
por: Li, Xiang, et al.
Publicado: (2026)
por: Li, Xiang, et al.
Publicado: (2026)
Offline Two-Player Zero-Sum Markov Games with KL Regularization
por: Chen, Claire, et al.
Publicado: (2026)
por: Chen, Claire, et al.
Publicado: (2026)
Disentangling Policy from Offline Task Representation Learning via Adversarial Data Augmentation
por: Jia, Chengxing, et al.
Publicado: (2024)
por: Jia, Chengxing, et al.
Publicado: (2024)
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only
por: Xiao, Wei, et al.
Publicado: (2025)
por: Xiao, Wei, et al.
Publicado: (2025)
Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy Search
por: Mou, Zhiyu, et al.
Publicado: (2025)
por: Mou, Zhiyu, et al.
Publicado: (2025)
Off-Policy Evaluation of Ranking Policies via Embedding-Space User Behavior Modeling
por: Takahashi, Tatsuki, et al.
Publicado: (2025)
por: Takahashi, Tatsuki, et al.
Publicado: (2025)
Policy-regularized Offline Multi-objective Reinforcement Learning
por: Lin, Qian, et al.
Publicado: (2024)
por: Lin, Qian, et al.
Publicado: (2024)
Dataset Clustering for Improved Offline Policy Learning
por: Wang, Qiang, et al.
Publicado: (2024)
por: Wang, Qiang, et al.
Publicado: (2024)
Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
por: Luo, Yu, et al.
Publicado: (2024)
por: Luo, Yu, et al.
Publicado: (2024)
Diffusion Policies with Value-Conditional Optimization for Offline Reinforcement Learning
por: Ma, Yunchang, et al.
Publicado: (2025)
por: Ma, Yunchang, et al.
Publicado: (2025)
MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics
por: Liu, Xinyu, et al.
Publicado: (2026)
por: Liu, Xinyu, et al.
Publicado: (2026)
Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
por: Zhai, Yuanzhao, et al.
Publicado: (2024)
por: Zhai, Yuanzhao, et al.
Publicado: (2024)
Active Reinforcement Learning Strategies for Offline Policy Improvement
por: Dukkipati, Ambedkar, et al.
Publicado: (2024)
por: Dukkipati, Ambedkar, et al.
Publicado: (2024)
Hypercube Policy Regularization Framework for Offline Reinforcement Learning
por: Shen, Yi, et al.
Publicado: (2024)
por: Shen, Yi, et al.
Publicado: (2024)
Ejemplares similares
-
Efficient Multi-Policy Evaluation for Reinforcement Learning
por: Liu, Shuze Daniel, et al.
Publicado: (2024) -
Efficient Policy Evaluation with Safety Constraint for Reinforcement Learning
por: Chen, Claire, et al.
Publicado: (2024) -
Doubly Optimal Policy Evaluation for Reinforcement Learning
por: Liu, Shuze Daniel, et al.
Publicado: (2024) -
The ODE Method for Stochastic Approximation and Reinforcement Learning with Markovian Noise
por: Liu, Shuze Daniel, et al.
Publicado: (2024) -
Convergence of Two-Timescale Markovian Stochastic Approximations with Applications in Reinforcement Learning
por: Mahadevan, Vagul, et al.
Publicado: (2026)