Joint Pricing and Resource Allocation: An Optimal Online-Learning Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Jianyu, Wang, Xuan, Wang, Yu-Xiang, Jiang, Jiashuo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918029525254144
author Xu, Jianyu
Wang, Xuan
Wang, Yu-Xiang
Jiang, Jiashuo
author_facet Xu, Jianyu
Wang, Xuan
Wang, Yu-Xiang
Jiang, Jiashuo
contents We study an online learning problem on dynamic pricing and resource allocation, where we make joint pricing and inventory decisions to maximize the overall net profit. We consider the stochastic dependence of demands on the price, which complicates the resource allocation process and introduces significant non-convexity and non-smoothness to the problem. To solve this problem, we develop an efficient algorithm that utilizes a "Lower-Confidence Bound (LCB)" meta-strategy over multiple OCO agents. Our algorithm achieves $\tilde{O}(\sqrt{Tmn})$ regret (for $m$ suppliers and $n$ consumers), which is optimal with respect to the time horizon $T$. Our results illustrate an effective integration of statistical learning methodologies with complex operations research problems.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18049
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Joint Pricing and Resource Allocation: An Optimal Online-Learning Approach
Xu, Jianyu
Wang, Xuan
Wang, Yu-Xiang
Jiang, Jiashuo
Machine Learning
Optimization and Control
91B06, 90B22, 91B24, 90B50, 90B80, 62P20
I.2.6
We study an online learning problem on dynamic pricing and resource allocation, where we make joint pricing and inventory decisions to maximize the overall net profit. We consider the stochastic dependence of demands on the price, which complicates the resource allocation process and introduces significant non-convexity and non-smoothness to the problem. To solve this problem, we develop an efficient algorithm that utilizes a "Lower-Confidence Bound (LCB)" meta-strategy over multiple OCO agents. Our algorithm achieves $\tilde{O}(\sqrt{Tmn})$ regret (for $m$ suppliers and $n$ consumers), which is optimal with respect to the time horizon $T$. Our results illustrate an effective integration of statistical learning methodologies with complex operations research problems.
title Joint Pricing and Resource Allocation: An Optimal Online-Learning Approach
topic Machine Learning
Optimization and Control
91B06, 90B22, 91B24, 90B50, 90B80, 62P20
I.2.6
url https://arxiv.org/abs/2501.18049