EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915712686096384 |
|---|---|
| author | Shen, Ye Pei, Dun Guo, Yiqiu Wang, Junying Guo, Yijin Zhang, Zicheng Jia, Qi Zhou, Jun Zhai, Guangtao |
| author_facet | Shen, Ye Pei, Dun Guo, Yiqiu Wang, Junying Guo, Yijin Zhang, Zicheng Jia, Qi Zhou, Jun Zhai, Guangtao |
| contents | Despite recent advances in understanding and leveraging long-range conversational memory, existing benchmarks still lack systematic evaluation of large language models(LLMs) across diverse memory dimensions, particularly in multi-session settings. In this work, we propose EvolMem, a new benchmark for assessing multi-session memory capabilities of LLMs and agent systems. EvolMem is grounded in cognitive psychology and encompasses both declarative and non-declarative memory, further decomposed into multiple fine-grained abilities. To construct the benchmark, we introduce a hybrid data synthesis framework that consists of topic-initiated generation and narrative-inspired transformations. This framework enables scalable generation of multi-session conversations with controllable complexity, accompanied by sample-specific evaluation guidelines. Extensive evaluation reveals that no LLM consistently outperforms others across all memory dimensions. Moreover, agent memory mechanisms do not necessarily enhance LLMs' capabilities and often exhibit notable efficiency limitations. Data and code will be released at https://github.com/shenye7436/EvolMem. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_03543 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory Shen, Ye Pei, Dun Guo, Yiqiu Wang, Junying Guo, Yijin Zhang, Zicheng Jia, Qi Zhou, Jun Zhai, Guangtao Computation and Language Despite recent advances in understanding and leveraging long-range conversational memory, existing benchmarks still lack systematic evaluation of large language models(LLMs) across diverse memory dimensions, particularly in multi-session settings. In this work, we propose EvolMem, a new benchmark for assessing multi-session memory capabilities of LLMs and agent systems. EvolMem is grounded in cognitive psychology and encompasses both declarative and non-declarative memory, further decomposed into multiple fine-grained abilities. To construct the benchmark, we introduce a hybrid data synthesis framework that consists of topic-initiated generation and narrative-inspired transformations. This framework enables scalable generation of multi-session conversations with controllable complexity, accompanied by sample-specific evaluation guidelines. Extensive evaluation reveals that no LLM consistently outperforms others across all memory dimensions. Moreover, agent memory mechanisms do not necessarily enhance LLMs' capabilities and often exhibit notable efficiency limitations. Data and code will be released at https://github.com/shenye7436/EvolMem. |
| title | EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2601.03543 |