Asymptotic properties of a multicolored random reinforced urn model with an application to multi-armed bandits

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Li, Hu, Jiang, Li, Jianghao, Bai, Zhidong
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910489053757440
author Yang, Li
Hu, Jiang
Li, Jianghao
Bai, Zhidong
author_facet Yang, Li
Hu, Jiang
Li, Jianghao
Bai, Zhidong
contents The random self-reinforcement mechanism, characterized by the principle of ``the rich get richer'', has demonstrated significant utility across various domains. One prominent model embodying this mechanism is the random reinforcement urn model. This paper investigates a multicolored, multiple-drawing variant of the random reinforced urn model. We establish the limiting behavior of the normalized urn composition and demonstrate strong convergence upon scaling the counts of each color. Additionally, we derive strong convergence estimators for the reinforcement means, i.e., for the expectations of the replacement matrix's diagonal elements, and prove their joint asymptotic normality. It is noteworthy that the estimators of the largest reinforcement mean are asymptotically independent of the estimators of the other smaller reinforcement means. Additionally, if a reinforcement mean is not the largest, the estimators of these smaller reinforcement means will also demonstrate asymptotic independence among themselves. Furthermore, we explore the parallels between the reinforced mechanisms in random reinforced urn models and multi-armed bandits, addressing hypothesis testing for expected payoffs in the latter context.
format Preprint
id arxiv_https___arxiv_org_abs_2406_10854
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Asymptotic properties of a multicolored random reinforced urn model with an application to multi-armed bandits
Yang, Li
Hu, Jiang
Li, Jianghao
Bai, Zhidong
Statistics Theory
The random self-reinforcement mechanism, characterized by the principle of ``the rich get richer'', has demonstrated significant utility across various domains. One prominent model embodying this mechanism is the random reinforcement urn model. This paper investigates a multicolored, multiple-drawing variant of the random reinforced urn model. We establish the limiting behavior of the normalized urn composition and demonstrate strong convergence upon scaling the counts of each color. Additionally, we derive strong convergence estimators for the reinforcement means, i.e., for the expectations of the replacement matrix's diagonal elements, and prove their joint asymptotic normality. It is noteworthy that the estimators of the largest reinforcement mean are asymptotically independent of the estimators of the other smaller reinforcement means. Additionally, if a reinforcement mean is not the largest, the estimators of these smaller reinforcement means will also demonstrate asymptotic independence among themselves. Furthermore, we explore the parallels between the reinforced mechanisms in random reinforced urn models and multi-armed bandits, addressing hypothesis testing for expected payoffs in the latter context.
title Asymptotic properties of a multicolored random reinforced urn model with an application to multi-armed bandits
topic Statistics Theory
url https://arxiv.org/abs/2406.10854