Solving multi-armed bandit problems using a chaotic microresonator comb

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cuevas, Jonathan, Iwami, Ryugo, Uchida, Atsushi, Minoshima, Kaoru, Kuse, Naoya
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908397881786368
author Cuevas, Jonathan
Iwami, Ryugo
Uchida, Atsushi
Minoshima, Kaoru
Kuse, Naoya
author_facet Cuevas, Jonathan
Iwami, Ryugo
Uchida, Atsushi
Minoshima, Kaoru
Kuse, Naoya
contents The Multi-Armed Bandit (MAB) problem, foundational to reinforcement learning-based decision-making, addresses the challenge of maximizing rewards amidst multiple uncertain choices. While algorithmic solutions are effective, their computational efficiency diminishes with increasing problem complexity. Photonic accelerators, leveraging temporal and spatial-temporal chaos, have emerged as promising alternatives. However, despite these advancements, current approaches either compromise computation speed or amplify system complexity. In this paper, we introduce a chaotic microresonator frequency comb (chaos comb) to tackle the MAB problem, where each comb mode is assigned to a slot machine. Through a proof-of-concept experiment, we employ 44 comb modes to address an MAB with 44 slot machines, demonstrating performance competitive with both conventional software algorithms and other photonic methods. Further, the scalability of decision making is explored with up to 512 slot machines using experimentally obtained temporal chaos in different time slots. Power-law scalability is achieved with an exponent of 0.96, outperforming conventional software-based algorithms. Moreover, we find that a numerically calculated chaos comb accurately reproduces experimental results, paving the way for discussions on strategies to increase the number of slot machines.
format Preprint
id arxiv_https___arxiv_org_abs_2308_10590
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Solving multi-armed bandit problems using a chaotic microresonator comb
Cuevas, Jonathan
Iwami, Ryugo
Uchida, Atsushi
Minoshima, Kaoru
Kuse, Naoya
Optics
Emerging Technologies
The Multi-Armed Bandit (MAB) problem, foundational to reinforcement learning-based decision-making, addresses the challenge of maximizing rewards amidst multiple uncertain choices. While algorithmic solutions are effective, their computational efficiency diminishes with increasing problem complexity. Photonic accelerators, leveraging temporal and spatial-temporal chaos, have emerged as promising alternatives. However, despite these advancements, current approaches either compromise computation speed or amplify system complexity. In this paper, we introduce a chaotic microresonator frequency comb (chaos comb) to tackle the MAB problem, where each comb mode is assigned to a slot machine. Through a proof-of-concept experiment, we employ 44 comb modes to address an MAB with 44 slot machines, demonstrating performance competitive with both conventional software algorithms and other photonic methods. Further, the scalability of decision making is explored with up to 512 slot machines using experimentally obtained temporal chaos in different time slots. Power-law scalability is achieved with an exponent of 0.96, outperforming conventional software-based algorithms. Moreover, we find that a numerically calculated chaos comb accurately reproduces experimental results, paving the way for discussions on strategies to increase the number of slot machines.
title Solving multi-armed bandit problems using a chaotic microresonator comb
topic Optics
Emerging Technologies
url https://arxiv.org/abs/2308.10590