Assessing Episodic Memory in LLMs with Sequence Order Recall Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pink, Mathis, Vo, Vy A., Wu, Qinyuan, Mu, Jianing, Turek, Javier S., Hasson, Uri, Norman, Kenneth A., Michelmann, Sebastian, Huth, Alexander, Toneva, Mariya
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913540378460160
author Pink, Mathis
Vo, Vy A.
Wu, Qinyuan
Mu, Jianing
Turek, Javier S.
Hasson, Uri
Norman, Kenneth A.
Michelmann, Sebastian
Huth, Alexander
Toneva, Mariya
author_facet Pink, Mathis
Vo, Vy A.
Wu, Qinyuan
Mu, Jianing
Turek, Javier S.
Hasson, Uri
Norman, Kenneth A.
Michelmann, Sebastian
Huth, Alexander
Toneva, Mariya
contents Current LLM benchmarks focus on evaluating models' memory of facts and semantic relations, primarily assessing semantic aspects of long-term memory. However, in humans, long-term memory also includes episodic memory, which links memories to their contexts, such as the time and place they occurred. The ability to contextualize memories is crucial for many cognitive tasks and everyday functions. This form of memory has not been evaluated in LLMs with existing benchmarks. To address the gap in evaluating memory in LLMs, we introduce Sequence Order Recall Tasks (SORT), which we adapt from tasks used to study episodic memory in cognitive psychology. SORT requires LLMs to recall the correct order of text segments, and provides a general framework that is both easily extendable and does not require any additional annotations. We present an initial evaluation dataset, Book-SORT, comprising 36k pairs of segments extracted from 9 books recently added to the public domain. Based on a human experiment with 155 participants, we show that humans can recall sequence order based on long-term memory of a book. We find that models can perform the task with high accuracy when relevant text is given in-context during the SORT evaluation. However, when presented with the book text only during training, LLMs' performance on SORT falls short. By allowing to evaluate more aspects of memory, we believe that SORT will aid in the emerging development of memory-augmented models.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08133
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Assessing Episodic Memory in LLMs with Sequence Order Recall Tasks
Pink, Mathis
Vo, Vy A.
Wu, Qinyuan
Mu, Jianing
Turek, Javier S.
Hasson, Uri
Norman, Kenneth A.
Michelmann, Sebastian
Huth, Alexander
Toneva, Mariya
Computation and Language
Artificial Intelligence
Machine Learning
Current LLM benchmarks focus on evaluating models' memory of facts and semantic relations, primarily assessing semantic aspects of long-term memory. However, in humans, long-term memory also includes episodic memory, which links memories to their contexts, such as the time and place they occurred. The ability to contextualize memories is crucial for many cognitive tasks and everyday functions. This form of memory has not been evaluated in LLMs with existing benchmarks. To address the gap in evaluating memory in LLMs, we introduce Sequence Order Recall Tasks (SORT), which we adapt from tasks used to study episodic memory in cognitive psychology. SORT requires LLMs to recall the correct order of text segments, and provides a general framework that is both easily extendable and does not require any additional annotations. We present an initial evaluation dataset, Book-SORT, comprising 36k pairs of segments extracted from 9 books recently added to the public domain. Based on a human experiment with 155 participants, we show that humans can recall sequence order based on long-term memory of a book. We find that models can perform the task with high accuracy when relevant text is given in-context during the SORT evaluation. However, when presented with the book text only during training, LLMs' performance on SORT falls short. By allowing to evaluate more aspects of memory, we believe that SORT will aid in the emerging development of memory-augmented models.
title Assessing Episodic Memory in LLMs with Sequence Order Recall Tasks
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.08133