Balancing Immediate Revenue and Future Off-Policy Evaluation in Coupon Allocation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Nishimura, Naoki, Kobayashi, Ken, Nakata, Kazuhide
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916385931657216
author Nishimura, Naoki
Kobayashi, Ken
Nakata, Kazuhide
author_facet Nishimura, Naoki
Kobayashi, Ken
Nakata, Kazuhide
contents Coupon allocation drives customer purchases and boosts revenue. However, it presents a fundamental trade-off between exploiting the current optimal policy to maximize immediate revenue and exploring alternative policies to collect data for future policy improvement via off-policy evaluation (OPE). To balance this trade-off, we propose a novel approach that combines a model-based revenue maximization policy and a randomized exploration policy for data collection. Our framework enables flexible adjustment of the mixture ratio between these two policies to optimize the balance between short-term revenue and future policy improvement. We formulate the problem of determining the optimal mixture ratio as multi-objective optimization, enabling quantitative evaluation of this trade-off. We empirically verified the effectiveness of the proposed mixed policy using synthetic data. Our main contributions are: (1) Demonstrating a mixed policy combining deterministic and probabilistic policies, flexibly adjusting the data collection vs. revenue trade-off. (2) Formulating the optimal mixture ratio problem as multi-objective optimization, enabling quantitative evaluation of this trade-off.
format Preprint
id arxiv_https___arxiv_org_abs_2407_11039
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Balancing Immediate Revenue and Future Off-Policy Evaluation in Coupon Allocation
Nishimura, Naoki
Kobayashi, Ken
Nakata, Kazuhide
Machine Learning
Artificial Intelligence
Coupon allocation drives customer purchases and boosts revenue. However, it presents a fundamental trade-off between exploiting the current optimal policy to maximize immediate revenue and exploring alternative policies to collect data for future policy improvement via off-policy evaluation (OPE). To balance this trade-off, we propose a novel approach that combines a model-based revenue maximization policy and a randomized exploration policy for data collection. Our framework enables flexible adjustment of the mixture ratio between these two policies to optimize the balance between short-term revenue and future policy improvement. We formulate the problem of determining the optimal mixture ratio as multi-objective optimization, enabling quantitative evaluation of this trade-off. We empirically verified the effectiveness of the proposed mixed policy using synthetic data. Our main contributions are: (1) Demonstrating a mixed policy combining deterministic and probabilistic policies, flexibly adjusting the data collection vs. revenue trade-off. (2) Formulating the optimal mixture ratio problem as multi-objective optimization, enabling quantitative evaluation of this trade-off.
title Balancing Immediate Revenue and Future Off-Policy Evaluation in Coupon Allocation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2407.11039