Model Selection for Off-policy Evaluation: New Algorithms and Experimental Protocol
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Pai, Zhao, Lingfeng, Agarwal, Shivangi, Liu, Jinghan, Huang, Audrey, Amortila, Philip, Jiang, Nan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Unifying View of Coverage in Linear Off-Policy Evaluation
by: Amortila, Philip, et al.
Published: (2026)
by: Amortila, Philip, et al.
Published: (2026)
Reinforcement Learning under Latent Dynamics: Toward Statistical and Algorithmic Modularity
by: Amortila, Philip, et al.
Published: (2024)
by: Amortila, Philip, et al.
Published: (2024)
Statistical Tractability of Off-policy Evaluation of History-dependent Policies in POMDPs
by: Zhang, Yuheng, et al.
Published: (2025)
by: Zhang, Yuheng, et al.
Published: (2025)
On the Curses of Future and History in Future-dependent Value Functions for Off-policy Evaluation
by: Zhang, Yuheng, et al.
Published: (2024)
by: Zhang, Yuheng, et al.
Published: (2024)
Learning Off-policy with Model-based Intrinsic Motivation For Active Online Exploration
by: Wang, Yibo, et al.
Published: (2024)
by: Wang, Yibo, et al.
Published: (2024)
Low Variance Off-policy Evaluation with State-based Importance Sampling
by: Bossens, David M., et al.
Published: (2022)
by: Bossens, David M., et al.
Published: (2022)
Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
ExO-PPO: an Extended Off-policy Proximal Policy Optimization Algorithm
by: Wang, Hanyong, et al.
Published: (2026)
by: Wang, Hanyong, et al.
Published: (2026)
Bootstrap Off-policy with World Model
by: Zhan, Guojian, et al.
Published: (2025)
by: Zhan, Guojian, et al.
Published: (2025)
Beyond the Parameters: A Technical Survey of Contextual Enrichment in Large Language Models: From In-Context Prompting to Causal Retrieval-Augmented Generation
by: Bansal, Prakhar, et al.
Published: (2026)
by: Bansal, Prakhar, et al.
Published: (2026)
Primal-Dual Spectral Representation for Off-policy Evaluation
by: Hu, Yang, et al.
Published: (2024)
by: Hu, Yang, et al.
Published: (2024)
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
by: Huang, Audrey, et al.
Published: (2025)
by: Huang, Audrey, et al.
Published: (2025)
Causal Deepsets for Off-policy Evaluation under Spatial or Spatio-temporal Interferences
by: Dai, Runpeng, et al.
Published: (2024)
by: Dai, Runpeng, et al.
Published: (2024)
Selecting Belief-State Approximations in Simulators with Latent States
by: Jiang, Nan
Published: (2025)
by: Jiang, Nan
Published: (2025)
Adapting Critic Match Loss Landscape Visualization to Off-policy Reinforcement Learning
by: Liu, Jingyi, et al.
Published: (2026)
by: Liu, Jingyi, et al.
Published: (2026)
Off-Policy Selection for Initiating Human-Centric Experimental Design
by: Gao, Ge, et al.
Published: (2024)
by: Gao, Ge, et al.
Published: (2024)
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
by: Liang, Jing, et al.
Published: (2025)
by: Liang, Jing, et al.
Published: (2025)
Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees
by: Jiang, Nan, et al.
Published: (2025)
by: Jiang, Nan, et al.
Published: (2025)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
by: Noukhovitch, Michael, et al.
Published: (2024)
by: Noukhovitch, Michael, et al.
Published: (2024)
Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation
by: Chaudhari, Shreyas, et al.
Published: (2024)
by: Chaudhari, Shreyas, et al.
Published: (2024)
MGAS: Multi-Granularity Architecture Search for Trade-Off Between Model Effectiveness and Efficiency
by: Liu, Xiaoyun, et al.
Published: (2023)
by: Liu, Xiaoyun, et al.
Published: (2023)
MSEval: A Dataset for Material Selection in Conceptual Design to Evaluate Algorithmic Models
by: Jain, Yash Patawari, et al.
Published: (2024)
by: Jain, Yash Patawari, et al.
Published: (2024)
A Note on Loss Functions and Error Compounding in Model-based Reinforcement Learning
by: Jiang, Nan
Published: (2024)
by: Jiang, Nan
Published: (2024)
Clustering Context in Off-Policy Evaluation
by: Guzman-Olivares, Daniel, et al.
Published: (2025)
by: Guzman-Olivares, Daniel, et al.
Published: (2025)
Concept-driven Off Policy Evaluation
by: Majumdar, Ritam, et al.
Published: (2024)
by: Majumdar, Ritam, et al.
Published: (2024)
Quotient DAGs for Off-Policy Evaluation:Forward-Flow Importance Sampling and Exact Slate Propensities
by: Xie, Ziwen, et al.
Published: (2026)
by: Xie, Ziwen, et al.
Published: (2026)
Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function Approximation
by: Che, Fengdi, et al.
Published: (2024)
by: Che, Fengdi, et al.
Published: (2024)
Tree-OPO: Off-policy Monte Carlo Tree-Guided Advantage Optimization for Multistep Reasoning
by: Huang, Bingning, et al.
Published: (2025)
by: Huang, Bingning, et al.
Published: (2025)
Pluralistic Off-policy Evaluation and Alignment
by: Huang, Chengkai, et al.
Published: (2025)
by: Huang, Chengkai, et al.
Published: (2025)
Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
by: Heuillet, Maxime, et al.
Published: (2025)
by: Heuillet, Maxime, et al.
Published: (2025)
Evaluating Supervised Machine Learning Models: Principles, Pitfalls, and Metric Selection
by: Liu, Xuanyan, et al.
Published: (2026)
by: Liu, Xuanyan, et al.
Published: (2026)
Automated Off-Policy Estimator Selection via Supervised Learning
by: Felicioni, Nicolò, et al.
Published: (2024)
by: Felicioni, Nicolò, et al.
Published: (2024)
A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
by: Patterson, Andrew, et al.
Published: (2021)
by: Patterson, Andrew, et al.
Published: (2021)
Learning Action Embeddings for Off-Policy Evaluation
by: Cief, Matej, et al.
Published: (2023)
by: Cief, Matej, et al.
Published: (2023)
Goal-oriented Transmission Scheduling: Structure-guided DRL with a Unified Dual On-policy and Off-policy Approach
by: Chen, Jiazheng, et al.
Published: (2025)
by: Chen, Jiazheng, et al.
Published: (2025)
MOBODY: Model Based Off-Dynamics Offline Reinforcement Learning
by: Guo, Yihong, et al.
Published: (2025)
by: Guo, Yihong, et al.
Published: (2025)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
by: Griffin, Charlie, et al.
Published: (2024)
by: Griffin, Charlie, et al.
Published: (2024)
SIMU: Selective Influence Machine Unlearning
by: Agarwal, Anu, et al.
Published: (2025)
by: Agarwal, Anu, et al.
Published: (2025)
The Path Not Taken: RLVR Provably Learns Off the Principals
by: Zhu, Hanqing, et al.
Published: (2025)
by: Zhu, Hanqing, et al.
Published: (2025)
Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights
by: Huang, Tzu-Heng, et al.
Published: (2025)
by: Huang, Tzu-Heng, et al.
Published: (2025)
Similar Items
-
A Unifying View of Coverage in Linear Off-Policy Evaluation
by: Amortila, Philip, et al.
Published: (2026) -
Reinforcement Learning under Latent Dynamics: Toward Statistical and Algorithmic Modularity
by: Amortila, Philip, et al.
Published: (2024) -
Statistical Tractability of Off-policy Evaluation of History-dependent Policies in POMDPs
by: Zhang, Yuheng, et al.
Published: (2025) -
On the Curses of Future and History in Future-dependent Value Functions for Off-policy Evaluation
by: Zhang, Yuheng, et al.
Published: (2024) -
Learning Off-policy with Model-based Intrinsic Motivation For Active Online Exploration
by: Wang, Yibo, et al.
Published: (2024)