Policy Learning for Balancing Short-Term and Long-Term Rewards

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Peng, Shen, Ziyu, Xie, Feng, Wang, Zhongyao, Liu, Chunchen, Zeng, Yan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910605007388672
author Wu, Peng
Shen, Ziyu
Xie, Feng
Wang, Zhongyao
Liu, Chunchen
Zeng, Yan
author_facet Wu, Peng
Shen, Ziyu
Xie, Feng
Wang, Zhongyao
Liu, Chunchen
Zeng, Yan
contents Empirical researchers and decision-makers spanning various domains frequently seek profound insights into the long-term impacts of interventions. While the significance of long-term outcomes is undeniable, an overemphasis on them may inadvertently overshadow short-term gains. Motivated by this, this paper formalizes a new framework for learning the optimal policy that effectively balances both long-term and short-term rewards, where some long-term outcomes are allowed to be missing. In particular, we first present the identifiability of both rewards under mild assumptions. Next, we deduce the semiparametric efficiency bounds, along with the consistency and asymptotic normality of their estimators. We also reveal that short-term outcomes, if associated, contribute to improving the estimator of the long-term reward. Based on the proposed estimators, we develop a principled policy learning approach and further derive the convergence rates of regret and estimation errors associated with the learned policy. Extensive experiments are conducted to validate the effectiveness of the proposed method, demonstrating its practical applicability.
format Preprint
id arxiv_https___arxiv_org_abs_2405_03329
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Policy Learning for Balancing Short-Term and Long-Term Rewards
Wu, Peng
Shen, Ziyu
Xie, Feng
Wang, Zhongyao
Liu, Chunchen
Zeng, Yan
Machine Learning
Empirical researchers and decision-makers spanning various domains frequently seek profound insights into the long-term impacts of interventions. While the significance of long-term outcomes is undeniable, an overemphasis on them may inadvertently overshadow short-term gains. Motivated by this, this paper formalizes a new framework for learning the optimal policy that effectively balances both long-term and short-term rewards, where some long-term outcomes are allowed to be missing. In particular, we first present the identifiability of both rewards under mild assumptions. Next, we deduce the semiparametric efficiency bounds, along with the consistency and asymptotic normality of their estimators. We also reveal that short-term outcomes, if associated, contribute to improving the estimator of the long-term reward. Based on the proposed estimators, we develop a principled policy learning approach and further derive the convergence rates of regret and estimation errors associated with the learned policy. Extensive experiments are conducted to validate the effectiveness of the proposed method, demonstrating its practical applicability.
title Policy Learning for Balancing Short-Term and Long-Term Rewards
topic Machine Learning
url https://arxiv.org/abs/2405.03329