Explore Reinforced: Equilibrium Approximation with Reinforcement Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910725303173120 |
|---|---|
| author | Yu, Ryan Nowak, Mateusz Xie, Qintong Feng, Michelle Yilin Chin, Peter |
| author_facet | Yu, Ryan Nowak, Mateusz Xie, Qintong Feng, Michelle Yilin Chin, Peter |
| contents | Current approximate Coarse Correlated Equilibria (CCE) algorithms struggle with equilibrium approximation for games in large stochastic environments but are theoretically guaranteed to converge to a strong solution concept. In contrast, modern Reinforcement Learning (RL) algorithms provide faster training yet yield weaker solutions. We introduce Exp3-IXrl - a blend of RL and game-theoretic approach, separating the RL agent's action selection from the equilibrium computation while preserving the integrity of the learning process. We demonstrate that our algorithm expands the application of equilibrium approximation algorithms to new environments. Specifically, we show the improved performance in a complex and adversarial cybersecurity network environment - the Cyber Operations Research Gym - and in the classical multi-armed bandit settings. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_02016 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Explore Reinforced: Equilibrium Approximation with Reinforcement Learning Yu, Ryan Nowak, Mateusz Xie, Qintong Feng, Michelle Yilin Chin, Peter Machine Learning Artificial Intelligence Computer Science and Game Theory Current approximate Coarse Correlated Equilibria (CCE) algorithms struggle with equilibrium approximation for games in large stochastic environments but are theoretically guaranteed to converge to a strong solution concept. In contrast, modern Reinforcement Learning (RL) algorithms provide faster training yet yield weaker solutions. We introduce Exp3-IXrl - a blend of RL and game-theoretic approach, separating the RL agent's action selection from the equilibrium computation while preserving the integrity of the learning process. We demonstrate that our algorithm expands the application of equilibrium approximation algorithms to new environments. Specifically, we show the improved performance in a complex and adversarial cybersecurity network environment - the Cyber Operations Research Gym - and in the classical multi-armed bandit settings. |
| title | Explore Reinforced: Equilibrium Approximation with Reinforcement Learning |
| topic | Machine Learning Artificial Intelligence Computer Science and Game Theory |
| url | https://arxiv.org/abs/2412.02016 |