Evaluating Long-Horizon Memory for Multi-Party Collaborative Dialogues

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hu, Chuanrui, Li, Tong, Gao, Xingze, Chen, Hongda, Bai, Yi, Xu, Dannong, Lin, Tianwei, Li, Xiaohong, Han, Yunyun, Pei, Jian, Deng, Yafeng
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912961009811456
author Hu, Chuanrui
Li, Tong
Gao, Xingze
Chen, Hongda
Bai, Yi
Xu, Dannong
Lin, Tianwei
Li, Xiaohong
Han, Yunyun
Pei, Jian
Deng, Yafeng
author_facet Hu, Chuanrui
Li, Tong
Gao, Xingze
Chen, Hongda
Bai, Yi
Xu, Dannong
Lin, Tianwei
Li, Xiaohong
Han, Yunyun
Pei, Jian
Deng, Yafeng
contents Long-term conversational memory in practical LLM applications is inherently collaborative: information is produced by multiple participants, scattered across groups and channels, revised over time, and implicitly grounded in roles and social context. Yet there is currently no established benchmark that evaluates memory under interaction patterns resembling real-world deployment, as existing benchmarks largely focus on dyadic or single-topic dialogues. In this paper, we introduce EverMemBench, the first benchmark designed for long-horizon collaborative memory, built from multi-party, multi-group conversations spanning over one million tokens with dense cross-topic interleaving, temporally evolving decisions, and role-conditioned personas. EverMemBench evaluates memory systems using 2400 QA pairs across three dimensions essential for real applications: fine-grained recall, memory awareness, and user profile understanding. Our evaluation reveals fundamental limitations of current systems: multi-hop reasoning collapses under multi-party attribution even with oracle evidence (26% accuracy), temporal reasoning fails without explicit version semantics beyond timestamps, and memory awareness is bottlenecked by retrieval, as similarity-based methods miss implicitly relevant information. EverMemBench thus represents a concrete step toward realistic evaluation of LLM memory and a cornerstone benchmark for developing next-generation LLMs that reason over time, roles, and collaborative interaction structure. Our benchmark and code are publicly available at https://github.com/EverMind-AI/EverMemBench.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01313
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evaluating Long-Horizon Memory for Multi-Party Collaborative Dialogues
Hu, Chuanrui
Li, Tong
Gao, Xingze
Chen, Hongda
Bai, Yi
Xu, Dannong
Lin, Tianwei
Li, Xiaohong
Han, Yunyun
Pei, Jian
Deng, Yafeng
Computation and Language
Artificial Intelligence
Long-term conversational memory in practical LLM applications is inherently collaborative: information is produced by multiple participants, scattered across groups and channels, revised over time, and implicitly grounded in roles and social context. Yet there is currently no established benchmark that evaluates memory under interaction patterns resembling real-world deployment, as existing benchmarks largely focus on dyadic or single-topic dialogues. In this paper, we introduce EverMemBench, the first benchmark designed for long-horizon collaborative memory, built from multi-party, multi-group conversations spanning over one million tokens with dense cross-topic interleaving, temporally evolving decisions, and role-conditioned personas. EverMemBench evaluates memory systems using 2400 QA pairs across three dimensions essential for real applications: fine-grained recall, memory awareness, and user profile understanding. Our evaluation reveals fundamental limitations of current systems: multi-hop reasoning collapses under multi-party attribution even with oracle evidence (26% accuracy), temporal reasoning fails without explicit version semantics beyond timestamps, and memory awareness is bottlenecked by retrieval, as similarity-based methods miss implicitly relevant information. EverMemBench thus represents a concrete step toward realistic evaluation of LLM memory and a cornerstone benchmark for developing next-generation LLMs that reason over time, roles, and collaborative interaction structure. Our benchmark and code are publicly available at https://github.com/EverMind-AI/EverMemBench.
title Evaluating Long-Horizon Memory for Multi-Party Collaborative Dialogues
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2602.01313