Do Large Language Models Learn Human-Like Strategic Preferences?
Fuente:
arXiv
Salvato in:
| Autori principali: | , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917792904642560 |
|---|---|
| author | Roberts, Jesse Moore, Kyle Fisher, Doug |
| author_facet | Roberts, Jesse Moore, Kyle Fisher, Doug |
| contents | In this paper, we evaluate whether LLMs learn to make human-like preference judgements in strategic scenarios as compared with known empirical results. Solar and Mistral are shown to exhibit stable value-based preference consistent with humans and exhibit human-like preference for cooperation in the prisoner's dilemma (including stake-size effect) and traveler's dilemma (including penalty-size effect). We establish a relationship between model size, value-based preference, and superficiality. Finally, results here show that models tending to be less brittle have relied on sliding window attention suggesting a potential link. Additionally, we contribute a novel method for constructing preference relations from arbitrary LLMs and support for a hypothesis regarding human behavior in the traveler's dilemma. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_08710 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Do Large Language Models Learn Human-Like Strategic Preferences? Roberts, Jesse Moore, Kyle Fisher, Doug Computer Science and Game Theory Artificial Intelligence In this paper, we evaluate whether LLMs learn to make human-like preference judgements in strategic scenarios as compared with known empirical results. Solar and Mistral are shown to exhibit stable value-based preference consistent with humans and exhibit human-like preference for cooperation in the prisoner's dilemma (including stake-size effect) and traveler's dilemma (including penalty-size effect). We establish a relationship between model size, value-based preference, and superficiality. Finally, results here show that models tending to be less brittle have relied on sliding window attention suggesting a potential link. Additionally, we contribute a novel method for constructing preference relations from arbitrary LLMs and support for a hypothesis regarding human behavior in the traveler's dilemma. |
| title | Do Large Language Models Learn Human-Like Strategic Preferences? |
| topic | Computer Science and Game Theory Artificial Intelligence |
| url | https://arxiv.org/abs/2404.08710 |