A Tale of Two Cities: Pessimism and Opportunism in Offline Dynamic Pricing
Fuente:
arXiv
Saved in:
| Main Authors: | Bian, Zeyu, Wang, Lan, Qi, Zhengling |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Double Fairness Policy Learning: Integrating Action Fairness and Outcome Fairness in Decision-making
by: Bian, Zeyu, et al.
Published: (2026)
by: Bian, Zeyu, et al.
Published: (2026)
Off-policy Evaluation in Doubly Inhomogeneous Environments
by: Bian, Zeyu, et al.
Published: (2023)
by: Bian, Zeyu, et al.
Published: (2023)
Offline Dynamic Inventory and Pricing Strategy: Addressing Censored and Dependent Demand
by: Gundem, Korel, et al.
Published: (2025)
by: Gundem, Korel, et al.
Published: (2025)
Beyond Demand Estimation: Consumer Surplus Evaluation via Cumulative Propensity Weights
by: Bian, Zeyu, et al.
Published: (2026)
by: Bian, Zeyu, et al.
Published: (2026)
From Robotics to Sepsis Treatment: Offline RL via Geometric Pessimism
by: Wanjari, Sarthak
Published: (2026)
by: Wanjari, Sarthak
Published: (2026)
Beyond Pessimism: Offline Learning in KL-regularized Games
by: Zhang, Yuheng, et al.
Published: (2026)
by: Zhang, Yuheng, et al.
Published: (2026)
Escaping Offline Pessimism: Vector-Field Reward Shaping for Safe Frontier Exploration
by: Roknilamouki, Amirhossein, et al.
Published: (2026)
by: Roknilamouki, Amirhossein, et al.
Published: (2026)
InSPO: Unlocking Intrinsic Self-Reflection for LLM Preference Optimization
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement Learning
by: Zhang, Dake, et al.
Published: (2024)
by: Zhang, Dake, et al.
Published: (2024)
PASTA: A Unified Framework for Offline Assortment Learning
by: Dong, Juncheng, et al.
Published: (2025)
by: Dong, Juncheng, et al.
Published: (2025)
Robust Offline Reinforcement learning with Heavy-Tailed Rewards
by: Zhu, Jin, et al.
Published: (2023)
by: Zhu, Jin, et al.
Published: (2023)
Unified PAC-Bayesian Study of Pessimism for Offline Policy Learning with Regularized Importance Sampling
by: Aouali, Imad, et al.
Published: (2024)
by: Aouali, Imad, et al.
Published: (2024)
Pessimism-Free Offline Learning in General-Sum Games via KL Regularization
by: Chen, Claire, et al.
Published: (2026)
by: Chen, Claire, et al.
Published: (2026)
OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
The Virtues of Pessimism in Inverse Reinforcement Learning
by: Wu, David, et al.
Published: (2024)
by: Wu, David, et al.
Published: (2024)
Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision Processes
by: Lu, Miao, et al.
Published: (2022)
by: Lu, Miao, et al.
Published: (2022)
Learning from Random Demonstrations: Offline Reinforcement Learning with Importance-Sampled Diffusion Models
by: Fang, Zeyu, et al.
Published: (2024)
by: Fang, Zeyu, et al.
Published: (2024)
Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning
by: Yu, Shuguang, et al.
Published: (2024)
by: Yu, Shuguang, et al.
Published: (2024)
DROP: Distributional and Regular Optimism and Pessimism for Reinforcement Learning
by: Kobayashi, Taisuke
Published: (2024)
by: Kobayashi, Taisuke
Published: (2024)
STEEL: Singularity-aware Reinforcement Learning
by: Chen, Xiaohong, et al.
Published: (2023)
by: Chen, Xiaohong, et al.
Published: (2023)
A Principled Path to Fitted Distributional Evaluation
by: Hong, Sungee, et al.
Published: (2025)
by: Hong, Sungee, et al.
Published: (2025)
Mitigating Preference Hacking in Policy Optimization with Pessimism
by: Gupta, Dhawal, et al.
Published: (2025)
by: Gupta, Dhawal, et al.
Published: (2025)
Manifold-Constrained Energy-Based Transition Models for Offline Reinforcement Learning
by: Fang, Zeyu, et al.
Published: (2026)
by: Fang, Zeyu, et al.
Published: (2026)
POLAR: A Pessimistic Model-based Policy Learning Algorithm for Dynamic Treatment Regimes
by: Zhang, Ruijia, et al.
Published: (2025)
by: Zhang, Ruijia, et al.
Published: (2025)
Distributional Off-policy Evaluation with Bellman Residual Minimization
by: Hong, Sungee, et al.
Published: (2024)
by: Hong, Sungee, et al.
Published: (2024)
From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism
by: Yu, Zhuohao, et al.
Published: (2026)
by: Yu, Zhuohao, et al.
Published: (2026)
Contextual Online Pricing with (Biased) Offline Data
by: Zhang, Yixuan, et al.
Published: (2025)
by: Zhang, Yixuan, et al.
Published: (2025)
Policy learning "without" overlap: Pessimism and generalized empirical Bernstein's inequality
by: Jin, Ying, et al.
Published: (2022)
by: Jin, Ying, et al.
Published: (2022)
Best-of-Tails: Bridging Optimism and Pessimism in Inference-Time Alignment
by: Hsu, Hsiang, et al.
Published: (2026)
by: Hsu, Hsiang, et al.
Published: (2026)
A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent
by: Xie, Shuo, et al.
Published: (2025)
by: Xie, Shuo, et al.
Published: (2025)
Pessimism Principle Can Be Effective: Towards a Framework for Zero-Shot Transfer Reinforcement Learning
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Mildly Constrained Evaluation Policy for Offline Reinforcement Learning
by: Xu, Linjie, et al.
Published: (2023)
by: Xu, Linjie, et al.
Published: (2023)
Quantile-Optimal Policy Learning under Unmeasured Confounding
by: Chen, Zhongren, et al.
Published: (2025)
by: Chen, Zhongren, et al.
Published: (2025)
Actor-Critic or Critic-Actor? A Tale of Two Time Scales
by: Bhatnagar, Shalabh, et al.
Published: (2022)
by: Bhatnagar, Shalabh, et al.
Published: (2022)
Incremental Sequence Labeling: A Tale of Two Shifts
by: Qiu, Shengjie, et al.
Published: (2024)
by: Qiu, Shengjie, et al.
Published: (2024)
A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning
by: Zhan, Qishi, et al.
Published: (2026)
by: Zhan, Qishi, et al.
Published: (2026)
Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning
by: Wang, Qi, et al.
Published: (2023)
by: Wang, Qi, et al.
Published: (2023)
A Tale of Two Temperatures: Simple, Efficient, and Diverse Sampling from Diffusion Language Models
by: Olausson, Theo X., et al.
Published: (2026)
by: Olausson, Theo X., et al.
Published: (2026)
A Tale of Two Learning Algorithms: Multiple Stream Random Walk and Asynchronous Gossip
by: Gholami, Peyman, et al.
Published: (2025)
by: Gholami, Peyman, et al.
Published: (2025)
Improved Regret Bound for Safe Reinforcement Learning via Tighter Cost Pessimism and Reward Optimism
by: Yu, Kihyun, et al.
Published: (2024)
by: Yu, Kihyun, et al.
Published: (2024)
Similar Items
-
Double Fairness Policy Learning: Integrating Action Fairness and Outcome Fairness in Decision-making
by: Bian, Zeyu, et al.
Published: (2026) -
Off-policy Evaluation in Doubly Inhomogeneous Environments
by: Bian, Zeyu, et al.
Published: (2023) -
Offline Dynamic Inventory and Pricing Strategy: Addressing Censored and Dependent Demand
by: Gundem, Korel, et al.
Published: (2025) -
Beyond Demand Estimation: Consumer Surplus Evaluation via Cumulative Propensity Weights
by: Bian, Zeyu, et al.
Published: (2026) -
From Robotics to Sepsis Treatment: Offline RL via Geometric Pessimism
by: Wanjari, Sarthak
Published: (2026)