Model-Based Reinforcement Learning with Double Oracle Efficiency in Policy Optimization and Offline Estimation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Haichen, Qian, Jian, Simchi-Levi, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Offline Oracle-Efficient Learning for Contextual MDPs via Layerwise Exploration-Exploitation Tradeoff
von: Qian, Jian, et al.
Veröffentlicht: (2024)
von: Qian, Jian, et al.
Veröffentlicht: (2024)
Perturbing the Derivative: Wild Refitting for Model-Free Evaluation of Machine Learning Models under Bregman Losses
von: Hu, Haichen, et al.
Veröffentlicht: (2025)
von: Hu, Haichen, et al.
Veröffentlicht: (2025)
Perturbing the Derivative: Doubly Wild Refitting for Model-Free Evaluation of Opaque Machine Learning Predictors
von: Hu, Haichen, et al.
Veröffentlicht: (2025)
von: Hu, Haichen, et al.
Veröffentlicht: (2025)
Pre-Trained AI Model Assisted Online Decision-Making under Missing Covariates: A Theoretical Perspective
von: Hu, Haichen, et al.
Veröffentlicht: (2025)
von: Hu, Haichen, et al.
Veröffentlicht: (2025)
Interleaved Resampling and Refitting: Data and Compute-Efficient Evaluation of Black-Box Predictors
von: Hu, Haichen, et al.
Veröffentlicht: (2026)
von: Hu, Haichen, et al.
Veröffentlicht: (2026)
Constrained Online Decision-Making: A Unified Framework
von: Hu, Haichen, et al.
Veröffentlicht: (2025)
von: Hu, Haichen, et al.
Veröffentlicht: (2025)
Contextual Online Decision Making with Infinite-Dimensional Functional Regression
von: Hu, Haichen, et al.
Veröffentlicht: (2025)
von: Hu, Haichen, et al.
Veröffentlicht: (2025)
Optimal Adaptive Experimental Design for Estimating Treatment Effect
von: Li, Jiachun, et al.
Veröffentlicht: (2024)
von: Li, Jiachun, et al.
Veröffentlicht: (2024)
A Simple and Optimal Policy Design with Safety against Heavy-Tailed Risk for Stochastic Bandits
von: Simchi-Levi, David, et al.
Veröffentlicht: (2022)
von: Simchi-Levi, David, et al.
Veröffentlicht: (2022)
OptiRepair: Closed-Loop Diagnosis and Repair of Supply Chain Optimization Models with LLM Agents
von: Ao, Ruicheng, et al.
Veröffentlicht: (2026)
von: Ao, Ruicheng, et al.
Veröffentlicht: (2026)
From Confounding to Learning: Dynamic Service Fee Pricing on Third-Party Platforms
von: Ai, Rui, et al.
Veröffentlicht: (2025)
von: Ai, Rui, et al.
Veröffentlicht: (2025)
Improving the Estimation of Lifetime Effects in A/B Testing via Treatment Locality
von: Chen, Shuze, et al.
Veröffentlicht: (2024)
von: Chen, Shuze, et al.
Veröffentlicht: (2024)
Policy-Based Trajectory Clustering in Offline Reinforcement Learning
von: Hu, Hao, et al.
Veröffentlicht: (2025)
von: Hu, Hao, et al.
Veröffentlicht: (2025)
Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning
von: Hu, Jifeng, et al.
Veröffentlicht: (2025)
von: Hu, Jifeng, et al.
Veröffentlicht: (2025)
Learning to Price with Resource Constraints: From Full Information to Machine-Learned Prices
von: Ao, Ruicheng, et al.
Veröffentlicht: (2025)
von: Ao, Ruicheng, et al.
Veröffentlicht: (2025)
Sobolev Norm Learning Rates for Conditional Mean Embeddings
von: Talwai, Prem, et al.
Veröffentlicht: (2021)
von: Talwai, Prem, et al.
Veröffentlicht: (2021)
Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization
von: Zhan, Simon Sinong, et al.
Veröffentlicht: (2025)
von: Zhan, Simon Sinong, et al.
Veröffentlicht: (2025)
Partial Identification under Missing Data Using Weak Shadow Variables from Pretrained Models
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
On the Reliability Limits of LLM-Based Multi-Agent Planning
von: Ao, Ruicheng, et al.
Veröffentlicht: (2026)
von: Ao, Ruicheng, et al.
Veröffentlicht: (2026)
Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy Optimization
von: Liu, Zongkai, et al.
Veröffentlicht: (2024)
von: Liu, Zongkai, et al.
Veröffentlicht: (2024)
Policy-regularized Offline Multi-objective Reinforcement Learning
von: Lin, Qian, et al.
Veröffentlicht: (2024)
von: Lin, Qian, et al.
Veröffentlicht: (2024)
SUMO: Search-Based Uncertainty Estimation for Model-Based Offline Reinforcement Learning
von: Qiao, Zhongjian, et al.
Veröffentlicht: (2024)
von: Qiao, Zhongjian, et al.
Veröffentlicht: (2024)
Solve Smart, Not Often: Policy Learning for Costly MILP Re-solving
von: Ai, Rui, et al.
Veröffentlicht: (2025)
von: Ai, Rui, et al.
Veröffentlicht: (2025)
Beyond ATE: Multi-Criteria Design for A/B Testing
von: Li, Jiachun, et al.
Veröffentlicht: (2025)
von: Li, Jiachun, et al.
Veröffentlicht: (2025)
Prediction-Guided Active Experiments
von: Ao, Ruicheng, et al.
Veröffentlicht: (2024)
von: Ao, Ruicheng, et al.
Veröffentlicht: (2024)
Constrained Policy Optimization with Explicit Behavior Density for Offline Reinforcement Learning
von: Zhang, Jing, et al.
Veröffentlicht: (2023)
von: Zhang, Jing, et al.
Veröffentlicht: (2023)
Constrained Latent Action Policies for Model-Based Offline Reinforcement Learning
von: Alles, Marvin, et al.
Veröffentlicht: (2024)
von: Alles, Marvin, et al.
Veröffentlicht: (2024)
Optimization Solution Functions as Deterministic Policies for Offline Reinforcement Learning
von: Khattar, Vanshaj, et al.
Veröffentlicht: (2024)
von: Khattar, Vanshaj, et al.
Veröffentlicht: (2024)
Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning
von: Zhang, Tianle, et al.
Veröffentlicht: (2024)
von: Zhang, Tianle, et al.
Veröffentlicht: (2024)
Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning
von: Gao, Chen-Xiao, et al.
Veröffentlicht: (2025)
von: Gao, Chen-Xiao, et al.
Veröffentlicht: (2025)
Adversarial Policy Optimization for Offline Preference-based Reinforcement Learning
von: Kang, Hyungkyu, et al.
Veröffentlicht: (2025)
von: Kang, Hyungkyu, et al.
Veröffentlicht: (2025)
The Value of Information in Resource-Constrained Pricing
von: Ao, Ruicheng, et al.
Veröffentlicht: (2026)
von: Ao, Ruicheng, et al.
Veröffentlicht: (2026)
On the Optimal Regret of Locally Private Linear Contextual Bandit
von: Li, Jiachun, et al.
Veröffentlicht: (2024)
von: Li, Jiachun, et al.
Veröffentlicht: (2024)
Your Offline Policy is Not Trustworthy: Bilevel Reinforcement Learning for Sequential Portfolio Optimization
von: Yuan, Haochen, et al.
Veröffentlicht: (2025)
von: Yuan, Haochen, et al.
Veröffentlicht: (2025)
Diffusion Policies with Value-Conditional Optimization for Offline Reinforcement Learning
von: Ma, Yunchang, et al.
Veröffentlicht: (2025)
von: Ma, Yunchang, et al.
Veröffentlicht: (2025)
Offline Trajectory Optimization for Offline Reinforcement Learning
von: Zhao, Ziqi, et al.
Veröffentlicht: (2024)
von: Zhao, Ziqi, et al.
Veröffentlicht: (2024)
Regret Distribution in Stochastic Bandits: Optimal Trade-off between Expectation and Tail Risk
von: Simchi-Levi, David, et al.
Veröffentlicht: (2023)
von: Simchi-Levi, David, et al.
Veröffentlicht: (2023)
Privacy Preserving Adaptive Experiment Design
von: Li, Jiachun, et al.
Veröffentlicht: (2024)
von: Li, Jiachun, et al.
Veröffentlicht: (2024)
Oracle-Guided Masked Contrastive Reinforcement Learning for Visuomotor Policies
von: Zhang, Yuhang, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhang, et al.
Veröffentlicht: (2025)
Is Mamba Compatible with Trajectory Optimization in Offline Reinforcement Learning?
von: Dai, Yang, et al.
Veröffentlicht: (2024)
von: Dai, Yang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Offline Oracle-Efficient Learning for Contextual MDPs via Layerwise Exploration-Exploitation Tradeoff
von: Qian, Jian, et al.
Veröffentlicht: (2024) -
Perturbing the Derivative: Wild Refitting for Model-Free Evaluation of Machine Learning Models under Bregman Losses
von: Hu, Haichen, et al.
Veröffentlicht: (2025) -
Perturbing the Derivative: Doubly Wild Refitting for Model-Free Evaluation of Opaque Machine Learning Predictors
von: Hu, Haichen, et al.
Veröffentlicht: (2025) -
Pre-Trained AI Model Assisted Online Decision-Making under Missing Covariates: A Theoretical Perspective
von: Hu, Haichen, et al.
Veröffentlicht: (2025) -
Interleaved Resampling and Refitting: Data and Compute-Efficient Evaluation of Black-Box Predictors
von: Hu, Haichen, et al.
Veröffentlicht: (2026)