Reasoning Abilities of Large Language Models: In-Depth Analysis on the Abstraction and Reasoning Corpus

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Seungpil, Sim, Woochang, Shin, Donghyeon, Seo, Wongyu, Park, Jiwon, Lee, Seokki, Hwang, Sanha, Kim, Sejin, Kim, Sundong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929601438023680
author Lee, Seungpil
Sim, Woochang
Shin, Donghyeon
Seo, Wongyu
Park, Jiwon
Lee, Seokki
Hwang, Sanha
Kim, Sejin
Kim, Sundong
author_facet Lee, Seungpil
Sim, Woochang
Shin, Donghyeon
Seo, Wongyu
Park, Jiwon
Lee, Seokki
Hwang, Sanha
Kim, Sejin
Kim, Sundong
contents The existing methods for evaluating the inference abilities of Large Language Models (LLMs) have been predominantly results-centric, making it challenging to assess the inference process comprehensively. We introduce a novel approach using the Abstraction and Reasoning Corpus (ARC) benchmark to evaluate the inference and contextual understanding abilities of LLMs in a process-centric manner, focusing on three key components from the Language of Thought Hypothesis (LoTH): Logical Coherence, Compositionality, and Productivity. Our carefully designed experiments reveal that while LLMs demonstrate some inference capabilities, they still significantly lag behind human-level reasoning in these three aspects. The main contribution of this paper lies in introducing the LoTH perspective, which provides a method for evaluating the reasoning process that conventional results-oriented approaches fail to capture, thereby offering new insights into the development of human-level reasoning in artificial intelligence systems.
format Preprint
id arxiv_https___arxiv_org_abs_2403_11793
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reasoning Abilities of Large Language Models: In-Depth Analysis on the Abstraction and Reasoning Corpus
Lee, Seungpil
Sim, Woochang
Shin, Donghyeon
Seo, Wongyu
Park, Jiwon
Lee, Seokki
Hwang, Sanha
Kim, Sejin
Kim, Sundong
Computation and Language
Artificial Intelligence
Emerging Technologies
Symbolic Computation
The existing methods for evaluating the inference abilities of Large Language Models (LLMs) have been predominantly results-centric, making it challenging to assess the inference process comprehensively. We introduce a novel approach using the Abstraction and Reasoning Corpus (ARC) benchmark to evaluate the inference and contextual understanding abilities of LLMs in a process-centric manner, focusing on three key components from the Language of Thought Hypothesis (LoTH): Logical Coherence, Compositionality, and Productivity. Our carefully designed experiments reveal that while LLMs demonstrate some inference capabilities, they still significantly lag behind human-level reasoning in these three aspects. The main contribution of this paper lies in introducing the LoTH perspective, which provides a method for evaluating the reasoning process that conventional results-oriented approaches fail to capture, thereby offering new insights into the development of human-level reasoning in artificial intelligence systems.
title Reasoning Abilities of Large Language Models: In-Depth Analysis on the Abstraction and Reasoning Corpus
topic Computation and Language
Artificial Intelligence
Emerging Technologies
Symbolic Computation
url https://arxiv.org/abs/2403.11793