CogBench: a large language model walks into a psychology lab

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Coda-Forno, Julian, Binz, Marcel, Wang, Jane X., Schulz, Eric
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917600147013632
author Coda-Forno, Julian
Binz, Marcel
Wang, Jane X.
Schulz, Eric
author_facet Coda-Forno, Julian
Binz, Marcel
Wang, Jane X.
Schulz, Eric
contents Large language models (LLMs) have significantly advanced the field of artificial intelligence. Yet, evaluating them comprehensively remains challenging. We argue that this is partly due to the predominant focus on performance metrics in most benchmarks. This paper introduces CogBench, a benchmark that includes ten behavioral metrics derived from seven cognitive psychology experiments. This novel approach offers a toolkit for phenotyping LLMs' behavior. We apply CogBench to 35 LLMs, yielding a rich and diverse dataset. We analyze this data using statistical multilevel modeling techniques, accounting for the nested dependencies among fine-tuned versions of specific LLMs. Our study highlights the crucial role of model size and reinforcement learning from human feedback (RLHF) in improving performance and aligning with human behavior. Interestingly, we find that open-source models are less risk-prone than proprietary models and that fine-tuning on code does not necessarily enhance LLMs' behavior. Finally, we explore the effects of prompt-engineering techniques. We discover that chain-of-thought prompting improves probabilistic reasoning, while take-a-step-back prompting fosters model-based behaviors.
format Preprint
id arxiv_https___arxiv_org_abs_2402_18225
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CogBench: a large language model walks into a psychology lab
Coda-Forno, Julian
Binz, Marcel
Wang, Jane X.
Schulz, Eric
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) have significantly advanced the field of artificial intelligence. Yet, evaluating them comprehensively remains challenging. We argue that this is partly due to the predominant focus on performance metrics in most benchmarks. This paper introduces CogBench, a benchmark that includes ten behavioral metrics derived from seven cognitive psychology experiments. This novel approach offers a toolkit for phenotyping LLMs' behavior. We apply CogBench to 35 LLMs, yielding a rich and diverse dataset. We analyze this data using statistical multilevel modeling techniques, accounting for the nested dependencies among fine-tuned versions of specific LLMs. Our study highlights the crucial role of model size and reinforcement learning from human feedback (RLHF) in improving performance and aligning with human behavior. Interestingly, we find that open-source models are less risk-prone than proprietary models and that fine-tuning on code does not necessarily enhance LLMs' behavior. Finally, we explore the effects of prompt-engineering techniques. We discover that chain-of-thought prompting improves probabilistic reasoning, while take-a-step-back prompting fosters model-based behaviors.
title CogBench: a large language model walks into a psychology lab
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2402.18225