Learning from Delayed Feedback in Games via Extra Prediction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fujimoto, Yuma, Abe, Kenshi, Ariu, Kaito
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908634918682624
author Fujimoto, Yuma
Abe, Kenshi
Ariu, Kaito
author_facet Fujimoto, Yuma
Abe, Kenshi
Ariu, Kaito
contents This study raises and addresses the problem of time-delayed feedback in learning in games. Because learning in games assumes that multiple agents independently learn their strategies, a discrepancy in optimization often emerges among the agents. To overcome this discrepancy, the prediction of the future reward is incorporated into algorithms, typically known as Optimistic Follow-the-Regularized-Leader (OFTRL). However, the time delay in observing the past rewards hinders the prediction. Indeed, this study firstly proves that even a single-step delay worsens the performance of OFTRL from the aspects of social regret and convergence. This study proposes the weighted OFTRL (WOFTRL), where the prediction vector of the next reward in OFTRL is weighted $n$ times. We further capture an intuition that the optimistic weight cancels out this time delay. We prove that when the optimistic weight exceeds the time delay, our WOFTRL recovers the good performances that social regret is constant in general-sum normal-form games, and the strategies last-iterate converge to the Nash equilibrium in poly-matrix zero-sum games. The theoretical results are supported and strengthened by our experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22426
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning from Delayed Feedback in Games via Extra Prediction
Fujimoto, Yuma
Abe, Kenshi
Ariu, Kaito
Machine Learning
Computer Science and Game Theory
Multiagent Systems
Optimization and Control
This study raises and addresses the problem of time-delayed feedback in learning in games. Because learning in games assumes that multiple agents independently learn their strategies, a discrepancy in optimization often emerges among the agents. To overcome this discrepancy, the prediction of the future reward is incorporated into algorithms, typically known as Optimistic Follow-the-Regularized-Leader (OFTRL). However, the time delay in observing the past rewards hinders the prediction. Indeed, this study firstly proves that even a single-step delay worsens the performance of OFTRL from the aspects of social regret and convergence. This study proposes the weighted OFTRL (WOFTRL), where the prediction vector of the next reward in OFTRL is weighted $n$ times. We further capture an intuition that the optimistic weight cancels out this time delay. We prove that when the optimistic weight exceeds the time delay, our WOFTRL recovers the good performances that social regret is constant in general-sum normal-form games, and the strategies last-iterate converge to the Nash equilibrium in poly-matrix zero-sum games. The theoretical results are supported and strengthened by our experiments.
title Learning from Delayed Feedback in Games via Extra Prediction
topic Machine Learning
Computer Science and Game Theory
Multiagent Systems
Optimization and Control
url https://arxiv.org/abs/2509.22426