Large Language Models Are Not Strong Abstract Reasoners

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gendron, Gaël, Bao, Qiming, Witbrock, Michael, Dobbie, Gillian
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909059876126720
author Gendron, Gaël
Bao, Qiming
Witbrock, Michael
Dobbie, Gillian
author_facet Gendron, Gaël
Bao, Qiming
Witbrock, Michael
Dobbie, Gillian
contents Large Language Models have shown tremendous performance on a large variety of natural language processing tasks, ranging from text comprehension to common sense reasoning. However, the mechanisms responsible for this success remain opaque, and it is unclear whether LLMs can achieve human-like cognitive capabilities or whether these models are still fundamentally circumscribed. Abstract reasoning is a fundamental task for cognition, consisting of finding and applying a general pattern from few data. Evaluating deep neural architectures on this task could give insight into their potential limitations regarding reasoning and their broad generalisation abilities, yet this is currently an under-explored area. In this paper, we introduce a new benchmark for evaluating language models beyond memorization on abstract reasoning tasks. We perform extensive evaluations of state-of-the-art LLMs, showing that they currently achieve very limited performance in contrast with other natural language tasks, even when applying techniques that have been shown to improve performance on other NLP tasks. We argue that guiding LLM generation to follow causal paths could help improve the generalisation and reasoning abilities of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2305_19555
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Large Language Models Are Not Strong Abstract Reasoners
Gendron, Gaël
Bao, Qiming
Witbrock, Michael
Dobbie, Gillian
Computation and Language
Machine Learning
I.2.2; I.2.3; I.2.7; I.5.1
Large Language Models have shown tremendous performance on a large variety of natural language processing tasks, ranging from text comprehension to common sense reasoning. However, the mechanisms responsible for this success remain opaque, and it is unclear whether LLMs can achieve human-like cognitive capabilities or whether these models are still fundamentally circumscribed. Abstract reasoning is a fundamental task for cognition, consisting of finding and applying a general pattern from few data. Evaluating deep neural architectures on this task could give insight into their potential limitations regarding reasoning and their broad generalisation abilities, yet this is currently an under-explored area. In this paper, we introduce a new benchmark for evaluating language models beyond memorization on abstract reasoning tasks. We perform extensive evaluations of state-of-the-art LLMs, showing that they currently achieve very limited performance in contrast with other natural language tasks, even when applying techniques that have been shown to improve performance on other NLP tasks. We argue that guiding LLM generation to follow causal paths could help improve the generalisation and reasoning abilities of LLMs.
title Large Language Models Are Not Strong Abstract Reasoners
topic Computation and Language
Machine Learning
I.2.2; I.2.3; I.2.7; I.5.1
url https://arxiv.org/abs/2305.19555