Investigating the Matching Law in Transformer-Based Reinforcement Learning Agents

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autori principali: Karakoc, Y, Sharifi, E., Gagana, M.D., Sarayloo, Z
Natura: Recurso digital
Lingua:inglese
Pubblicazione: Zenodo 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866902085377720320
author Karakoc, Y
Sharifi, E.
Gagana, M.D.
Sarayloo, Z
author_facet Karakoc, Y
Sharifi, E.
Gagana, M.D.
Sarayloo, Z
contents <div> <div> <div> <div> <p>This study examines whether a Transformer-based reinforcement learning agent follows the Matching Law—the principle that organisms allocate responses in proportion to rewards. We trained a Transformer agent on a two-choice sequential decision task with varying reward contingencies to see if its choice behavior matches relative reward rates. The Transformer’s attention-based architecture was leveraged to learn temporal patterns and adapt its policy. We found that the agent’s choice proportions closely mirrored the reward ratios across conditions, demonstrating adherence to the Matching Law. Log-log analyses showed an approximately linear relationship between response and reward ratios. These results suggest that even advanced AI agents like Transformers can spontaneously exhibit Matching Law behavior, offering insights into links between artificial and biological learning.</p> </div> </div> </div> </div>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_15404010
institution Zenodo
language eng
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle Investigating the Matching Law in Transformer-Based Reinforcement Learning Agents
Karakoc, Y
Sharifi, E.
Gagana, M.D.
Sarayloo, Z
Matching Law
Reinforcement learning
Sequence Decision-Making
Transformer
<div> <div> <div> <div> <p>This study examines whether a Transformer-based reinforcement learning agent follows the Matching Law—the principle that organisms allocate responses in proportion to rewards. We trained a Transformer agent on a two-choice sequential decision task with varying reward contingencies to see if its choice behavior matches relative reward rates. The Transformer’s attention-based architecture was leveraged to learn temporal patterns and adapt its policy. We found that the agent’s choice proportions closely mirrored the reward ratios across conditions, demonstrating adherence to the Matching Law. Log-log analyses showed an approximately linear relationship between response and reward ratios. These results suggest that even advanced AI agents like Transformers can spontaneously exhibit Matching Law behavior, offering insights into links between artificial and biological learning.</p> </div> </div> </div> </div>
title Investigating the Matching Law in Transformer-Based Reinforcement Learning Agents
topic Matching Law
Reinforcement learning
Sequence Decision-Making
Transformer
url https://doi.org/10.5281/zenodo.15404010