CloneMem: Benchmarking Long-Term Memory for AI Clones

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Hu, Sen, Zhang, Zhiyu, Wei, Yuxiang, Han, Xueran, Tang, Zhenheng, Wang, Huacan, Chen, Ronghao
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915722461970432
author Hu, Sen
Zhang, Zhiyu
Wei, Yuxiang
Han, Xueran
Tang, Zhenheng
Wang, Huacan
Chen, Ronghao
author_facet Hu, Sen
Zhang, Zhiyu
Wei, Yuxiang
Han, Xueran
Tang, Zhenheng
Wang, Huacan
Chen, Ronghao
contents AI Clones aim to simulate an individual's thoughts and behaviors to enable long-term, personalized interaction, placing stringent demands on memory systems to model experiences, emotions, and opinions over time. Existing memory benchmarks primarily rely on user-agent conversational histories, which are temporally fragmented and insufficient for capturing continuous life trajectories. We introduce CloneMem, a benchmark for evaluating longterm memory in AI Clone scenarios grounded in non-conversational digital traces, including diaries, social media posts, and emails, spanning one to three years. CloneMem adopts a hierarchical data construction framework to ensure longitudinal coherence and defines tasks that assess an agent's ability to track evolving personal states. Experiments show that current memory mechanisms struggle in this setting, highlighting open challenges for life-grounded personalized AI. Code and dataset are available at https://github.com/AvatarMemory/CloneMemBench
format Preprint
id arxiv_https___arxiv_org_abs_2601_07023
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CloneMem: Benchmarking Long-Term Memory for AI Clones
Hu, Sen
Zhang, Zhiyu
Wei, Yuxiang
Han, Xueran
Tang, Zhenheng
Wang, Huacan
Chen, Ronghao
Artificial Intelligence
AI Clones aim to simulate an individual's thoughts and behaviors to enable long-term, personalized interaction, placing stringent demands on memory systems to model experiences, emotions, and opinions over time. Existing memory benchmarks primarily rely on user-agent conversational histories, which are temporally fragmented and insufficient for capturing continuous life trajectories. We introduce CloneMem, a benchmark for evaluating longterm memory in AI Clone scenarios grounded in non-conversational digital traces, including diaries, social media posts, and emails, spanning one to three years. CloneMem adopts a hierarchical data construction framework to ensure longitudinal coherence and defines tasks that assess an agent's ability to track evolving personal states. Experiments show that current memory mechanisms struggle in this setting, highlighting open challenges for life-grounded personalized AI. Code and dataset are available at https://github.com/AvatarMemory/CloneMemBench
title CloneMem: Benchmarking Long-Term Memory for AI Clones
topic Artificial Intelligence
url https://arxiv.org/abs/2601.07023