Analyzing Narrative Processing in Large Language Models (LLMs): Using GPT4 to test BERT

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Krauss, Patrick, Hösch, Jannik, Metzner, Claus, Maier, Andreas, Uhrig, Peter, Schilling, Achim
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929334785146880
author Krauss, Patrick
Hösch, Jannik
Metzner, Claus
Maier, Andreas
Uhrig, Peter
Schilling, Achim
author_facet Krauss, Patrick
Hösch, Jannik
Metzner, Claus
Maier, Andreas
Uhrig, Peter
Schilling, Achim
contents The ability to transmit and receive complex information via language is unique to humans and is the basis of traditions, culture and versatile social interactions. Through the disruptive introduction of transformer based large language models (LLMs) humans are not the only entity to "understand" and produce language any more. In the present study, we have performed the first steps to use LLMs as a model to understand fundamental mechanisms of language processing in neural networks, in order to make predictions and generate hypotheses on how the human brain does language processing. Thus, we have used ChatGPT to generate seven different stylistic variations of ten different narratives (Aesop's fables). We used these stories as input for the open source LLM BERT and have analyzed the activation patterns of the hidden units of BERT using multi-dimensional scaling and cluster analysis. We found that the activation vectors of the hidden units cluster according to stylistic variations in earlier layers of BERT (1) than narrative content (4-5). Despite the fact that BERT consists of 12 identical building blocks that are stacked and trained on large text corpora, the different layers perform different tasks. This is a very useful model of the human brain, where self-similar structures, i.e. different areas of the cerebral cortex, can have different functions and are therefore well suited to processing language in a very efficient way. The proposed approach has the potential to open the black box of LLMs on the one hand, and might be a further step to unravel the neural processes underlying human language processing and cognition in general.
format Preprint
id arxiv_https___arxiv_org_abs_2405_02024
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Analyzing Narrative Processing in Large Language Models (LLMs): Using GPT4 to test BERT
Krauss, Patrick
Hösch, Jannik
Metzner, Claus
Maier, Andreas
Uhrig, Peter
Schilling, Achim
Computation and Language
Artificial Intelligence
The ability to transmit and receive complex information via language is unique to humans and is the basis of traditions, culture and versatile social interactions. Through the disruptive introduction of transformer based large language models (LLMs) humans are not the only entity to "understand" and produce language any more. In the present study, we have performed the first steps to use LLMs as a model to understand fundamental mechanisms of language processing in neural networks, in order to make predictions and generate hypotheses on how the human brain does language processing. Thus, we have used ChatGPT to generate seven different stylistic variations of ten different narratives (Aesop's fables). We used these stories as input for the open source LLM BERT and have analyzed the activation patterns of the hidden units of BERT using multi-dimensional scaling and cluster analysis. We found that the activation vectors of the hidden units cluster according to stylistic variations in earlier layers of BERT (1) than narrative content (4-5). Despite the fact that BERT consists of 12 identical building blocks that are stacked and trained on large text corpora, the different layers perform different tasks. This is a very useful model of the human brain, where self-similar structures, i.e. different areas of the cerebral cortex, can have different functions and are therefore well suited to processing language in a very efficient way. The proposed approach has the potential to open the black box of LLMs on the one hand, and might be a further step to unravel the neural processes underlying human language processing and cognition in general.
title Analyzing Narrative Processing in Large Language Models (LLMs): Using GPT4 to test BERT
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2405.02024