EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Ye, Pei, Dun, Guo, Yiqiu, Wang, Junying, Guo, Yijin, Zhang, Zicheng, Jia, Qi, Zhou, Jun, Zhai, Guangtao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915712686096384
author Shen, Ye
Pei, Dun
Guo, Yiqiu
Wang, Junying
Guo, Yijin
Zhang, Zicheng
Jia, Qi
Zhou, Jun
Zhai, Guangtao
author_facet Shen, Ye
Pei, Dun
Guo, Yiqiu
Wang, Junying
Guo, Yijin
Zhang, Zicheng
Jia, Qi
Zhou, Jun
Zhai, Guangtao
contents Despite recent advances in understanding and leveraging long-range conversational memory, existing benchmarks still lack systematic evaluation of large language models(LLMs) across diverse memory dimensions, particularly in multi-session settings. In this work, we propose EvolMem, a new benchmark for assessing multi-session memory capabilities of LLMs and agent systems. EvolMem is grounded in cognitive psychology and encompasses both declarative and non-declarative memory, further decomposed into multiple fine-grained abilities. To construct the benchmark, we introduce a hybrid data synthesis framework that consists of topic-initiated generation and narrative-inspired transformations. This framework enables scalable generation of multi-session conversations with controllable complexity, accompanied by sample-specific evaluation guidelines. Extensive evaluation reveals that no LLM consistently outperforms others across all memory dimensions. Moreover, agent memory mechanisms do not necessarily enhance LLMs' capabilities and often exhibit notable efficiency limitations. Data and code will be released at https://github.com/shenye7436/EvolMem.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03543
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
Shen, Ye
Pei, Dun
Guo, Yiqiu
Wang, Junying
Guo, Yijin
Zhang, Zicheng
Jia, Qi
Zhou, Jun
Zhai, Guangtao
Computation and Language
Despite recent advances in understanding and leveraging long-range conversational memory, existing benchmarks still lack systematic evaluation of large language models(LLMs) across diverse memory dimensions, particularly in multi-session settings. In this work, we propose EvolMem, a new benchmark for assessing multi-session memory capabilities of LLMs and agent systems. EvolMem is grounded in cognitive psychology and encompasses both declarative and non-declarative memory, further decomposed into multiple fine-grained abilities. To construct the benchmark, we introduce a hybrid data synthesis framework that consists of topic-initiated generation and narrative-inspired transformations. This framework enables scalable generation of multi-session conversations with controllable complexity, accompanied by sample-specific evaluation guidelines. Extensive evaluation reveals that no LLM consistently outperforms others across all memory dimensions. Moreover, agent memory mechanisms do not necessarily enhance LLMs' capabilities and often exhibit notable efficiency limitations. Data and code will be released at https://github.com/shenye7436/EvolMem.
title EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
topic Computation and Language
url https://arxiv.org/abs/2601.03543