If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fan, Siqi, Huang, Xiusheng, Yao, Yiqun, Fang, Xuezhi, Liu, Kang, Han, Peng, Shang, Shuo, Sun, Aixin, Wang, Yequan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911587146661888
author Fan, Siqi
Huang, Xiusheng
Yao, Yiqun
Fang, Xuezhi
Liu, Kang
Han, Peng
Shang, Shuo
Sun, Aixin
Wang, Yequan
author_facet Fan, Siqi
Huang, Xiusheng
Yao, Yiqun
Fang, Xuezhi
Liu, Kang
Han, Peng
Shang, Shuo
Sun, Aixin
Wang, Yequan
contents Large language models (LLMs) can carry out human-like dialogue, but unlike humans, they are stateless due to the superposition property. However, during multi-turn, multi-agent interactions, LLMs begin to exhibit consistent, character-like behaviors, hinting at a form of emergent lifelong learning. Despite this, existing benchmarks often fail to capture these dynamics, primarily focusing on static, open-ended evaluations. To address this gap, we introduce LIFESTATE-BENCH, a benchmark designed to assess lifelong learning in LLMs. It features two episodic datasets: Hamlet and a synthetic script collection, rich in narrative structure and character interactions. Our fact checking evaluation probes models' self-awareness, episodic memory retrieval, and relationship tracking, across both parametric and non-parametric approaches. Experiments on models like Llama3.1-8B, GPT-4-turbo, and DeepSeek R1, we demonstrate that nonparametric methods significantly outperform parametric ones in managing stateful learning. However, all models exhibit challenges with catastrophic forgetting as interactions extend, highlighting the need for further advancements in lifelong learning.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23514
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMs
Fan, Siqi
Huang, Xiusheng
Yao, Yiqun
Fang, Xuezhi
Liu, Kang
Han, Peng
Shang, Shuo
Sun, Aixin
Wang, Yequan
Computation and Language
Artificial Intelligence
Large language models (LLMs) can carry out human-like dialogue, but unlike humans, they are stateless due to the superposition property. However, during multi-turn, multi-agent interactions, LLMs begin to exhibit consistent, character-like behaviors, hinting at a form of emergent lifelong learning. Despite this, existing benchmarks often fail to capture these dynamics, primarily focusing on static, open-ended evaluations. To address this gap, we introduce LIFESTATE-BENCH, a benchmark designed to assess lifelong learning in LLMs. It features two episodic datasets: Hamlet and a synthetic script collection, rich in narrative structure and character interactions. Our fact checking evaluation probes models' self-awareness, episodic memory retrieval, and relationship tracking, across both parametric and non-parametric approaches. Experiments on models like Llama3.1-8B, GPT-4-turbo, and DeepSeek R1, we demonstrate that nonparametric methods significantly outperform parametric ones in managing stateful learning. However, all models exhibit challenges with catastrophic forgetting as interactions extend, highlighting the need for further advancements in lifelong learning.
title If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.23514