Haunted House: A text-based game for comparing the flexibility of mental models in humans and LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Puppart, Brett, Paltmann, Paul-Henry, Aru, Jaan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908276198735872
author Puppart, Brett
Paltmann, Paul-Henry
Aru, Jaan
author_facet Puppart, Brett
Paltmann, Paul-Henry
Aru, Jaan
contents This study introduces "Haunted House" a novel text-based game designed to compare the performance of humans and large language models (LLMs) in model-based reasoning. Players must escape from a house containing nine rooms in a 3x3 grid layout while avoiding the ghost. They are guided by verbal clues that they get each time they move. In Study 1, the results from 98 human participants revealed a success rate of 31.6%, significantly outperforming seven state-of-the-art LLMs tested. Out of 140 attempts across seven LLMs, only one attempt resulted in a pass by Claude 3 Opus. Preliminary results suggested that GPT o3-mini-high performance might be higher, but not at the human level. Further analysis of 29 human participants' moves in Study 2 indicated that LLMs frequently struggled with random and illogical moves, while humans exhibited such errors less frequently. Our findings suggest that current LLMs encounter difficulties in tasks that demand active model-based reasoning, offering inspiration for future benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16437
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Haunted House: A text-based game for comparing the flexibility of mental models in humans and LLMs
Puppart, Brett
Paltmann, Paul-Henry
Aru, Jaan
Human-Computer Interaction
Artificial Intelligence
Neurons and Cognition
This study introduces "Haunted House" a novel text-based game designed to compare the performance of humans and large language models (LLMs) in model-based reasoning. Players must escape from a house containing nine rooms in a 3x3 grid layout while avoiding the ghost. They are guided by verbal clues that they get each time they move. In Study 1, the results from 98 human participants revealed a success rate of 31.6%, significantly outperforming seven state-of-the-art LLMs tested. Out of 140 attempts across seven LLMs, only one attempt resulted in a pass by Claude 3 Opus. Preliminary results suggested that GPT o3-mini-high performance might be higher, but not at the human level. Further analysis of 29 human participants' moves in Study 2 indicated that LLMs frequently struggled with random and illogical moves, while humans exhibited such errors less frequently. Our findings suggest that current LLMs encounter difficulties in tasks that demand active model-based reasoning, offering inspiration for future benchmarks.
title Haunted House: A text-based game for comparing the flexibility of mental models in humans and LLMs
topic Human-Computer Interaction
Artificial Intelligence
Neurons and Cognition
url https://arxiv.org/abs/2503.16437