On Passivity, Reinforcement Learning and Higher-Order Learning in Multi-Agent Finite Games
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2018
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914995190628352 |
|---|---|
| author | Gao, Bolin Pavel, Lacra |
| author_facet | Gao, Bolin Pavel, Lacra |
| contents | In this paper, we propose a passivity-based methodology for analysis and design of reinforcement learning in multi-agent finite games. Starting from a known exponentially-discounted reinforcement learning scheme, we show that convergence to a Nash distribution can be shown in the class of games characterized by the monotonicity property of their (negative) payoff. We further exploit passivity to propose a class of higher-order schemes that preserve convergence properties, can improve the speed of convergence and can even converge in cases whereby their first-order counterpart fail to converge. We demonstrate these properties through numerical simulations for several representative games. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_1808_04464 |
| institution | arXiv |
| publishDate | 2018 |
| record_format | arxiv |
| spellingShingle | On Passivity, Reinforcement Learning and Higher-Order Learning in Multi-Agent Finite Games Gao, Bolin Pavel, Lacra Optimization and Control Computer Science and Game Theory Systems and Control In this paper, we propose a passivity-based methodology for analysis and design of reinforcement learning in multi-agent finite games. Starting from a known exponentially-discounted reinforcement learning scheme, we show that convergence to a Nash distribution can be shown in the class of games characterized by the monotonicity property of their (negative) payoff. We further exploit passivity to propose a class of higher-order schemes that preserve convergence properties, can improve the speed of convergence and can even converge in cases whereby their first-order counterpart fail to converge. We demonstrate these properties through numerical simulations for several representative games. |
| title | On Passivity, Reinforcement Learning and Higher-Order Learning in Multi-Agent Finite Games |
| topic | Optimization and Control Computer Science and Game Theory Systems and Control |
| url | https://arxiv.org/abs/1808.04464 |