A Reduction-based Framework for Sequential Decision Making with Delayed Feedback

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Yunchang, Zhong, Han, Wu, Tianhao, Liu, Bin, Wang, Liwei, Du, Simon S.
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911789753565184
author Yang, Yunchang
Zhong, Han
Wu, Tianhao
Liu, Bin
Wang, Liwei
Du, Simon S.
author_facet Yang, Yunchang
Zhong, Han
Wu, Tianhao
Liu, Bin
Wang, Liwei
Du, Simon S.
contents We study stochastic delayed feedback in general multi-agent sequential decision making, which includes bandits, single-agent Markov decision processes (MDPs), and Markov games (MGs). We propose a novel reduction-based framework, which turns any multi-batched algorithm for sequential decision making with instantaneous feedback into a sample-efficient algorithm that can handle stochastic delays in sequential decision making. By plugging different multi-batched algorithms into our framework, we provide several examples demonstrating that our framework not only matches or improves existing results for bandits, tabular MDPs, and tabular MGs, but also provides the first line of studies on delays in sequential decision making with function approximation. In summary, we provide a complete set of sharp results for multi-agent sequential decision making with delayed feedback.
format Preprint
id arxiv_https___arxiv_org_abs_2302_01477
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A Reduction-based Framework for Sequential Decision Making with Delayed Feedback
Yang, Yunchang
Zhong, Han
Wu, Tianhao
Liu, Bin
Wang, Liwei
Du, Simon S.
Machine Learning
We study stochastic delayed feedback in general multi-agent sequential decision making, which includes bandits, single-agent Markov decision processes (MDPs), and Markov games (MGs). We propose a novel reduction-based framework, which turns any multi-batched algorithm for sequential decision making with instantaneous feedback into a sample-efficient algorithm that can handle stochastic delays in sequential decision making. By plugging different multi-batched algorithms into our framework, we provide several examples demonstrating that our framework not only matches or improves existing results for bandits, tabular MDPs, and tabular MGs, but also provides the first line of studies on delays in sequential decision making with function approximation. In summary, we provide a complete set of sharp results for multi-agent sequential decision making with delayed feedback.
title A Reduction-based Framework for Sequential Decision Making with Delayed Feedback
topic Machine Learning
url https://arxiv.org/abs/2302.01477