Staff View: :: Library Catalog

Saved in:

Bibliographic Details
Main Authors:	Li, Chao, Liu, Cailiang, Gao, Ang, Deng, Kexin, Zhang, Shu, Xu, Langping, Shi, Xiaotong, Ding, Xionghao, Pei, Jian, Jiang, Xun
Format:	Preprint
Published:	2026
Subjects:	Artificial Intelligence
Online Access:	https://arxiv.org/abs/2604.02834
Tags:	Add Tag No Tags, Be the first to tag this record!

_version_	1866908934739066880
author	Li, Chao Liu, Cailiang Gao, Ang Deng, Kexin Zhang, Shu Xu, Langping Shi, Xiaotong Ding, Xionghao Pei, Jian Jiang, Xun
author_facet	Li, Chao Liu, Cailiang Gao, Ang Deng, Kexin Zhang, Shu Xu, Langping Shi, Xiaotong Ding, Xionghao Pei, Jian Jiang, Xun
contents	Longitudinal health agents must reason across multi-source trajectories that combine continuous device streams, sparse clinical exams, and episodic life events - yet evaluating them is hard: real-world data cannot be released at scale, and temporally grounded attribution questions seldom admit definitive answers without structured ground truth. We present ESL-Bench, an event-driven synthesis framework and benchmark providing 100 synthetic users, each with a 1-5 year trajectory comprising a health profile, a multi-phase narrative plan, daily device measurements, periodic exam records, and an event log with explicit per-indicator impact parameters. Each indicator follows a baseline stochastic process driven by discrete events with sigmoid-onset, exponential-decay kernels under saturation and projection constraints; a hybrid pipeline delegates sparse semantic artifacts to LLM-based planning and dense indicator dynamics to algorithmic simulation with hard physiological bounds. Users are each paired with 100 evaluation queries across five dimensions - Lookup, Trend, Comparison, Anomaly, Explanation - stratified into Easy, Medium, and Hard tiers, with all ground-truth answers programmatically computable from the recorded event-indicator relationships. Evaluating 13 methods spanning LLMs with tools, DB-native agents, and memory-augmented RAG, we find that DB agents (48-58%) substantially outperform memory RAG baselines (30-38%), with the gap concentrated on Comparison and Explanation queries where multi-hop reasoning and evidence attribution are required.
format	Preprint
id	arxiv_https___arxiv_org_abs_2604_02834
institution	arXiv
publishDate	2026
record_format	arxiv
spellingShingle	ESL-Bench: An Event-Driven Synthetic Longitudinal Benchmark for Health Agents Li, Chao Liu, Cailiang Gao, Ang Deng, Kexin Zhang, Shu Xu, Langping Shi, Xiaotong Ding, Xionghao Pei, Jian Jiang, Xun Artificial Intelligence Longitudinal health agents must reason across multi-source trajectories that combine continuous device streams, sparse clinical exams, and episodic life events - yet evaluating them is hard: real-world data cannot be released at scale, and temporally grounded attribution questions seldom admit definitive answers without structured ground truth. We present ESL-Bench, an event-driven synthesis framework and benchmark providing 100 synthetic users, each with a 1-5 year trajectory comprising a health profile, a multi-phase narrative plan, daily device measurements, periodic exam records, and an event log with explicit per-indicator impact parameters. Each indicator follows a baseline stochastic process driven by discrete events with sigmoid-onset, exponential-decay kernels under saturation and projection constraints; a hybrid pipeline delegates sparse semantic artifacts to LLM-based planning and dense indicator dynamics to algorithmic simulation with hard physiological bounds. Users are each paired with 100 evaluation queries across five dimensions - Lookup, Trend, Comparison, Anomaly, Explanation - stratified into Easy, Medium, and Hard tiers, with all ground-truth answers programmatically computable from the recorded event-indicator relationships. Evaluating 13 methods spanning LLMs with tools, DB-native agents, and memory-augmented RAG, we find that DB agents (48-58%) substantially outperform memory RAG baselines (30-38%), with the gap concentrated on Comparison and Explanation queries where multi-hop reasoning and evidence attribution are required.
title	ESL-Bench: An Event-Driven Synthetic Longitudinal Benchmark for Health Agents
topic	Artificial Intelligence
url	https://arxiv.org/abs/2604.02834

Similar Items