Sample Complexity of Preference-Based Nonparametric Off-Policy Evaluation with Deep Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zihao, Ji, Xiang, Chen, Minshuo, Wang, Mengdi |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional Manifolds
by: Xu, Zhenghao, et al.
Published: (2023)
by: Xu, Zhenghao, et al.
Published: (2023)
Nonparametric Classification on Low Dimensional Manifolds using Overparameterized Convolutional Residual Networks
by: Zhang, Zixuan, et al.
Published: (2023)
by: Zhang, Zixuan, et al.
Published: (2023)
Theoretical Insights for Diffusion Guidance: A Case Study for Gaussian Mixture Models
by: Wu, Yuchen, et al.
Published: (2024)
by: Wu, Yuchen, et al.
Published: (2024)
Diffusion Model for Manifold Data: Score Decomposition, Curvature, and Statistical Complexity
by: Zhang, Zixuan, et al.
Published: (2026)
by: Zhang, Zixuan, et al.
Published: (2026)
Provable Statistical Rates for Consistency Diffusion Models
by: Dou, Zehao, et al.
Published: (2024)
by: Dou, Zehao, et al.
Published: (2024)
Unveil Conditional Diffusion Models with Classifier-free Guidance: A Sharp Statistical Theory
by: Fu, Hengyu, et al.
Published: (2024)
by: Fu, Hengyu, et al.
Published: (2024)
An Overview of Diffusion Models: Applications, Guided Generation, Statistical Rates and Optimization
by: Chen, Minshuo, et al.
Published: (2024)
by: Chen, Minshuo, et al.
Published: (2024)
From Weighting to Modeling: A Nonparametric Estimator for Off-Policy Evaluation
by: Zhu, Rong J. B.
Published: (2026)
by: Zhu, Rong J. B.
Published: (2026)
Diffusion Model for Data-Driven Black-Box Optimization
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
A Theoretical Perspective for Speculative Decoding Algorithm
by: Yin, Ming, et al.
Published: (2024)
by: Yin, Ming, et al.
Published: (2024)
Gradient Guidance for Diffusion Models: An Optimization Perspective
by: Guo, Yingqing, et al.
Published: (2024)
by: Guo, Yingqing, et al.
Published: (2024)
On the Role of Preference Variance in Preference Optimization
by: Guo, Jiacheng, et al.
Published: (2025)
by: Guo, Jiacheng, et al.
Published: (2025)
Diffusion Transformer Captures Spatial-Temporal Dependencies: A Theory for Gaussian Process Data
by: Fu, Hengyu, et al.
Published: (2024)
by: Fu, Hengyu, et al.
Published: (2024)
Nonparametric LLM Evaluation from Preference Data
by: Frauen, Dennis, et al.
Published: (2026)
by: Frauen, Dennis, et al.
Published: (2026)
Distributional Off-Policy Evaluation with Deep Quantile Process Regression
by: Kuang, Qi, et al.
Published: (2026)
by: Kuang, Qi, et al.
Published: (2026)
Quotient DAGs for Off-Policy Evaluation:Forward-Flow Importance Sampling and Exact Slate Propensities
by: Xie, Ziwen, et al.
Published: (2026)
by: Xie, Ziwen, et al.
Published: (2026)
Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
Diffusion Transformers for Imputation: Statistical Efficiency and Uncertainty Quantification
by: Ye, Zeqi, et al.
Published: (2025)
by: Ye, Zeqi, et al.
Published: (2025)
Double Duality: Variational Primal-Dual Policy Optimization for Constrained Reinforcement Learning
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
by: Lee, Haanvid, et al.
Published: (2024)
by: Lee, Haanvid, et al.
Published: (2024)
Self-Augmented Preference Optimization: Off-Policy Paradigms for Language Model Alignment
by: Yin, Yueqin, et al.
Published: (2024)
by: Yin, Yueqin, et al.
Published: (2024)
Regularized DeepIV with Model Selection
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
Deep Learning meets Nonparametric Regression: Are Weight-Decayed DNNs Locally Adaptive?
by: Zhang, Kaiqi, et al.
Published: (2022)
by: Zhang, Kaiqi, et al.
Published: (2022)
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning
by: Mao, Yixiu, et al.
Published: (2026)
by: Mao, Yixiu, et al.
Published: (2026)
On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures
by: Yin, Ming, et al.
Published: (2025)
by: Yin, Ming, et al.
Published: (2025)
Cross-Validated Off-Policy Evaluation
by: Cief, Matej, et al.
Published: (2024)
by: Cief, Matej, et al.
Published: (2024)
Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation
by: Zhou, Hongyi, et al.
Published: (2025)
by: Zhou, Hongyi, et al.
Published: (2025)
Off-Policy Evaluation Under Nonignorable Missing Data
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
Off-Policy Evaluation Using Information Borrowing and Context-Based Switching
by: Dasgupta, Sutanoy, et al.
Published: (2021)
by: Dasgupta, Sutanoy, et al.
Published: (2021)
Off-Policy Evaluation for Ranking Policies under Deterministic Logging Policies
by: Tanaka, Koichi, et al.
Published: (2026)
by: Tanaka, Koichi, et al.
Published: (2026)
Doubly-Robust Off-Policy Evaluation with Estimated Logging Policy
by: Lee, Kyungbok, et al.
Published: (2024)
by: Lee, Kyungbok, et al.
Published: (2024)
Long-term Off-Policy Evaluation and Learning
by: Saito, Yuta, et al.
Published: (2024)
by: Saito, Yuta, et al.
Published: (2024)
Off-Policy Value-Based Reinforcement Learning for Large Language Models
by: Wang, Peng-Yuan, et al.
Published: (2026)
by: Wang, Peng-Yuan, et al.
Published: (2026)
Inference for Deep Neural Network Estimators in Generalized Nonparametric Models
by: Meng, Xuran, et al.
Published: (2025)
by: Meng, Xuran, et al.
Published: (2025)
Compression Repair for Feedforward Neural Networks Based on Model Equivalence Evaluation
by: Mo, Zihao, et al.
Published: (2024)
by: Mo, Zihao, et al.
Published: (2024)
Sample-Efficient Policy Constraint Offline Deep Reinforcement Learning based on Sample Filtering
by: Chen, Yuanhao, et al.
Published: (2025)
by: Chen, Yuanhao, et al.
Published: (2025)
Off-Policy Evaluation of Slate Bandit Policies via Optimizing Abstraction
by: Kiyohara, Haruka, et al.
Published: (2024)
by: Kiyohara, Haruka, et al.
Published: (2024)
Clustering Context in Off-Policy Evaluation
by: Guzman-Olivares, Daniel, et al.
Published: (2025)
by: Guzman-Olivares, Daniel, et al.
Published: (2025)
Concept-driven Off Policy Evaluation
by: Majumdar, Ritam, et al.
Published: (2024)
by: Majumdar, Ritam, et al.
Published: (2024)
Theoretical Analysis of Inductive Biases in Deep Convolutional Networks
by: Wang, Zihao, et al.
Published: (2023)
by: Wang, Zihao, et al.
Published: (2023)
Similar Items
-
Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional Manifolds
by: Xu, Zhenghao, et al.
Published: (2023) -
Nonparametric Classification on Low Dimensional Manifolds using Overparameterized Convolutional Residual Networks
by: Zhang, Zixuan, et al.
Published: (2023) -
Theoretical Insights for Diffusion Guidance: A Case Study for Gaussian Mixture Models
by: Wu, Yuchen, et al.
Published: (2024) -
Diffusion Model for Manifold Data: Score Decomposition, Curvature, and Statistical Complexity
by: Zhang, Zixuan, et al.
Published: (2026) -
Provable Statistical Rates for Consistency Diffusion Models
by: Dou, Zehao, et al.
Published: (2024)