MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yihao, Xu, Haoran, Gu, Renjie, Ye, Yixuan, Chen, Xinyi, Mu, Xinyu, Gao, Yuan, Guo, Chunxiao, Wei, Peng, Gu, Jinjie, Li, Huan, Chen, Ke, Shou, Lidan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913116423454720
author Wang, Yihao
Xu, Haoran
Gu, Renjie
Ye, Yixuan
Chen, Xinyi
Mu, Xinyu
Gao, Yuan
Guo, Chunxiao
Wei, Peng
Gu, Jinjie
Li, Huan
Chen, Ke
Shou, Lidan
author_facet Wang, Yihao
Xu, Haoran
Gu, Renjie
Ye, Yixuan
Chen, Xinyi
Mu, Xinyu
Gao, Yuan
Guo, Chunxiao
Wei, Peng
Gu, Jinjie
Li, Huan
Chen, Ke
Shou, Lidan
contents The large-scale deployment of personalized healthcare agents demands memory mechanisms that are exceptionally precise, safe, and capable of long-term clinical tracking. However, existing benchmarks primarily focus on daily open-domain conversations, failing to capture the high-stakes complexity of real-world medical applications. Motivated by the stringent production requirements of an industry-leading health management agent serving tens of millions of active users, we introduce MedMemoryBench. We develop a human-agent collaborative pipeline to synthesize highly realistic, long-horizon medical trajectories based on clinically grounded, synthetic patient archetypes. This process yields a massive, expertly validated dataset comprising approximately 2,000 sessions and 16,000 interaction turns. Crucially, MedMemoryBench departs from traditional static evaluations by pioneering an "evaluate-while-constructing" streaming assessment protocol, which precisely mirrors dynamic memory accumulation in production environments. Furthermore, we formalize and systematically investigate the critical phenomenon of memory saturation, where sustained information influx actively degrades retrieval and reasoning robustness. Comprehensive benchmarking reveals severe bottlenecks in mainstream architectures, particularly concerning complex medical reasoning and noise resilience. By exposing these fundamental flaws, MedMemoryBench establishes a vital foundation for developing robust, production-ready medical agents.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11814
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
Wang, Yihao
Xu, Haoran
Gu, Renjie
Ye, Yixuan
Chen, Xinyi
Mu, Xinyu
Gao, Yuan
Guo, Chunxiao
Wei, Peng
Gu, Jinjie
Li, Huan
Chen, Ke
Shou, Lidan
Artificial Intelligence
68T07, 68T50
I.2.7; I.2.1
The large-scale deployment of personalized healthcare agents demands memory mechanisms that are exceptionally precise, safe, and capable of long-term clinical tracking. However, existing benchmarks primarily focus on daily open-domain conversations, failing to capture the high-stakes complexity of real-world medical applications. Motivated by the stringent production requirements of an industry-leading health management agent serving tens of millions of active users, we introduce MedMemoryBench. We develop a human-agent collaborative pipeline to synthesize highly realistic, long-horizon medical trajectories based on clinically grounded, synthetic patient archetypes. This process yields a massive, expertly validated dataset comprising approximately 2,000 sessions and 16,000 interaction turns. Crucially, MedMemoryBench departs from traditional static evaluations by pioneering an "evaluate-while-constructing" streaming assessment protocol, which precisely mirrors dynamic memory accumulation in production environments. Furthermore, we formalize and systematically investigate the critical phenomenon of memory saturation, where sustained information influx actively degrades retrieval and reasoning robustness. Comprehensive benchmarking reveals severe bottlenecks in mainstream architectures, particularly concerning complex medical reasoning and noise resilience. By exposing these fundamental flaws, MedMemoryBench establishes a vital foundation for developing robust, production-ready medical agents.
title MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
topic Artificial Intelligence
68T07, 68T50
I.2.7; I.2.1
url https://arxiv.org/abs/2605.11814