A Principled Path to Fitted Distributional Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hong, Sungee, Wang, Jiayi, Qi, Zhengling, Wong, Raymond K. W. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Distributional Off-policy Evaluation with Bellman Residual Minimization
von: Hong, Sungee, et al.
Veröffentlicht: (2024)
von: Hong, Sungee, et al.
Veröffentlicht: (2024)
A Fine-grained Analysis of Fitted Q-evaluation: Beyond Parametric Models
von: Wang, Jiayi, et al.
Veröffentlicht: (2024)
von: Wang, Jiayi, et al.
Veröffentlicht: (2024)
A Tale of Two Cities: Pessimism and Opportunism in Offline Dynamic Pricing
von: Bian, Zeyu, et al.
Veröffentlicht: (2024)
von: Bian, Zeyu, et al.
Veröffentlicht: (2024)
Off-policy Evaluation in Doubly Inhomogeneous Environments
von: Bian, Zeyu, et al.
Veröffentlicht: (2023)
von: Bian, Zeyu, et al.
Veröffentlicht: (2023)
Double Fairness Policy Learning: Integrating Action Fairness and Outcome Fairness in Decision-making
von: Bian, Zeyu, et al.
Veröffentlicht: (2026)
von: Bian, Zeyu, et al.
Veröffentlicht: (2026)
CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning
von: Sun, Ke, et al.
Veröffentlicht: (2026)
von: Sun, Ke, et al.
Veröffentlicht: (2026)
InSPO: Unlocking Intrinsic Self-Reflection for LLM Preference Optimization
von: Li, Yu, et al.
Veröffentlicht: (2025)
von: Li, Yu, et al.
Veröffentlicht: (2025)
STEEL: Singularity-aware Reinforcement Learning
von: Chen, Xiaohong, et al.
Veröffentlicht: (2023)
von: Chen, Xiaohong, et al.
Veröffentlicht: (2023)
Beyond Demand Estimation: Consumer Surplus Evaluation via Cumulative Propensity Weights
von: Bian, Zeyu, et al.
Veröffentlicht: (2026)
von: Bian, Zeyu, et al.
Veröffentlicht: (2026)
Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning
von: Yu, Shuguang, et al.
Veröffentlicht: (2024)
von: Yu, Shuguang, et al.
Veröffentlicht: (2024)
OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
Offline Dynamic Inventory and Pricing Strategy: Addressing Censored and Dependent Demand
von: Gundem, Korel, et al.
Veröffentlicht: (2025)
von: Gundem, Korel, et al.
Veröffentlicht: (2025)
Quantile-Optimal Policy Learning under Unmeasured Confounding
von: Chen, Zhongren, et al.
Veröffentlicht: (2025)
von: Chen, Zhongren, et al.
Veröffentlicht: (2025)
PASTA: A Unified Framework for Offline Assortment Learning
von: Dong, Juncheng, et al.
Veröffentlicht: (2025)
von: Dong, Juncheng, et al.
Veröffentlicht: (2025)
POLAR: A Pessimistic Model-based Policy Learning Algorithm for Dynamic Treatment Regimes
von: Zhang, Ruijia, et al.
Veröffentlicht: (2025)
von: Zhang, Ruijia, et al.
Veröffentlicht: (2025)
Robust Offline Reinforcement learning with Heavy-Tailed Rewards
von: Zhu, Jin, et al.
Veröffentlicht: (2023)
von: Zhu, Jin, et al.
Veröffentlicht: (2023)
Fair Regression under Demographic Parity: A Unified Framework
von: Feng, Yongzhen, et al.
Veröffentlicht: (2026)
von: Feng, Yongzhen, et al.
Veröffentlicht: (2026)
Beyond One-Size-Fits-All: Tailored Benchmarks for Efficient Evaluation
von: Yuan, Peiwen, et al.
Veröffentlicht: (2025)
von: Yuan, Peiwen, et al.
Veröffentlicht: (2025)
CP Degeneracy in Tensor Regression
von: Zhou, Ya, et al.
Veröffentlicht: (2020)
von: Zhou, Ya, et al.
Veröffentlicht: (2020)
Revisiting Neural Processes via Fourier Transform and Volterra Series
von: Mohseni, Peiman, et al.
Veröffentlicht: (2026)
von: Mohseni, Peiman, et al.
Veröffentlicht: (2026)
Boosting In-Context Learning in LLMs Through the Lens of Classical Supervised Learning
von: Gundem, Korel, et al.
Veröffentlicht: (2025)
von: Gundem, Korel, et al.
Veröffentlicht: (2025)
Learning Robust Treatment Rules for Censored Data
von: Cui, Yifan, et al.
Veröffentlicht: (2024)
von: Cui, Yifan, et al.
Veröffentlicht: (2024)
Sequential Knockoffs for Variable Selection in Reinforcement Learning
von: Ma, Tao, et al.
Veröffentlicht: (2023)
von: Ma, Tao, et al.
Veröffentlicht: (2023)
Lagrangian Flow Matching: A Least-Action Framework for Principled Path Design
von: Du, Shukai, et al.
Veröffentlicht: (2026)
von: Du, Shukai, et al.
Veröffentlicht: (2026)
Reinforcement Learning with Continuous Actions Under Unmeasured Confounding
von: Li, Yuhan, et al.
Veröffentlicht: (2025)
von: Li, Yuhan, et al.
Veröffentlicht: (2025)
Lepskii Principle for Distributed Kernel Ridge Regression
von: Lin, Shao-Bo
Veröffentlicht: (2024)
von: Lin, Shao-Bo
Veröffentlicht: (2024)
Principled Out-of-Distribution Generalization via Simplicity
von: Ge, Jiawei, et al.
Veröffentlicht: (2025)
von: Ge, Jiawei, et al.
Veröffentlicht: (2025)
Density-aware Sample-specific Attack
von: Wang, Qiyuan, et al.
Veröffentlicht: (2026)
von: Wang, Qiyuan, et al.
Veröffentlicht: (2026)
Distributional Off-Policy Evaluation with Deep Quantile Process Regression
von: Kuang, Qi, et al.
Veröffentlicht: (2026)
von: Kuang, Qi, et al.
Veröffentlicht: (2026)
Rethinking Test-time Likelihood: The Likelihood Path Principle and Its Application to OOD Detection
von: Huang, Sicong, et al.
Veröffentlicht: (2024)
von: Huang, Sicong, et al.
Veröffentlicht: (2024)
Affinity Is Not Enough: Recovering the Free Energy Principle in Mixture-of-Experts
von: Wong, Man Yung
Veröffentlicht: (2026)
von: Wong, Man Yung
Veröffentlicht: (2026)
FedFitTech: A Baseline in Federated Learning for Fitness Tracking
von: Oz, Zeyneddin, et al.
Veröffentlicht: (2025)
von: Oz, Zeyneddin, et al.
Veröffentlicht: (2025)
Personalized Path Recourse for Reinforcement Learning Agents
von: Hong, Dat, et al.
Veröffentlicht: (2023)
von: Hong, Dat, et al.
Veröffentlicht: (2023)
On the Empirical Power of Goodness-of-Fit Tests in Watermark Detection
von: He, Weiqing, et al.
Veröffentlicht: (2025)
von: He, Weiqing, et al.
Veröffentlicht: (2025)
Fitted $Q$ Evaluation Without Bellman Completeness via Stationary Weighting
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
Distribution Fitting for Combating Mode Collapse in Generative Adversarial Networks
von: Gong, Yanxiang, et al.
Veröffentlicht: (2022)
von: Gong, Yanxiang, et al.
Veröffentlicht: (2022)
Out-of-Distribution Detection for Continual Learning: Design Principles and Benchmarking
von: Gupta, Srishti, et al.
Veröffentlicht: (2025)
von: Gupta, Srishti, et al.
Veröffentlicht: (2025)
On the Usefulness of the Fit-on-the-Test View on Evaluating Calibration of Classifiers
von: Kängsepp, Markus, et al.
Veröffentlicht: (2022)
von: Kängsepp, Markus, et al.
Veröffentlicht: (2022)
Robust Multi-Agent Path Finding under Observation Attacks: A Principled Adversarial-Plus-Smoothing Training Recipe
von: Ahmed, Riad
Veröffentlicht: (2026)
von: Ahmed, Riad
Veröffentlicht: (2026)
Finite-Time Bounds for Average-Reward Fitted Q-Iteration
von: Lee, Jongmin, et al.
Veröffentlicht: (2025)
von: Lee, Jongmin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Distributional Off-policy Evaluation with Bellman Residual Minimization
von: Hong, Sungee, et al.
Veröffentlicht: (2024) -
A Fine-grained Analysis of Fitted Q-evaluation: Beyond Parametric Models
von: Wang, Jiayi, et al.
Veröffentlicht: (2024) -
A Tale of Two Cities: Pessimism and Opportunism in Offline Dynamic Pricing
von: Bian, Zeyu, et al.
Veröffentlicht: (2024) -
Off-policy Evaluation in Doubly Inhomogeneous Environments
von: Bian, Zeyu, et al.
Veröffentlicht: (2023) -
Double Fairness Policy Learning: Integrating Action Fairness and Outcome Fairness in Decision-making
von: Bian, Zeyu, et al.
Veröffentlicht: (2026)