ESL-Bench: An Event-Driven Synthetic Longitudinal Benchmark for Health Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Chao, Liu, Cailiang, Gao, Ang, Deng, Kexin, Zhang, Shu, Xu, Langping, Shi, Xiaotong, Ding, Xionghao, Pei, Jian, Jiang, Xun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FlightBench: Benchmarking Learning-based Methods for Ego-vision-based Quadrotors Navigation
von: Yu, Shu-Ang, et al.
Veröffentlicht: (2024)
von: Yu, Shu-Ang, et al.
Veröffentlicht: (2024)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
SWE-Bench 5G: Benchmarking AI Coding Agents on Telecom Network Engineering Tasks
von: Chen, Jiao, et al.
Veröffentlicht: (2026)
von: Chen, Jiao, et al.
Veröffentlicht: (2026)
HSCO-Bench: An Agent-Driven End-to-End Hardware-Software Co-design Benchmark for Systems-on-Chip
von: Tsai, Pei-Huan, et al.
Veröffentlicht: (2026)
von: Tsai, Pei-Huan, et al.
Veröffentlicht: (2026)
PERMA: Benchmarking Personalized Memory Agents via Event-Driven Preference and Realistic Task Environments
von: Liu, Shuochen, et al.
Veröffentlicht: (2026)
von: Liu, Shuochen, et al.
Veröffentlicht: (2026)
MedDialBench: Benchmarking LLM Diagnostic Robustness under Parametric Adversarial Patient Behaviors
von: Luo, Xiaotian, et al.
Veröffentlicht: (2026)
von: Luo, Xiaotian, et al.
Veröffentlicht: (2026)
LifeAgentBench: A Multi-dimensional Benchmark and Agent for Personal Health Assistants in Digital Health
von: Tian, Ye, et al.
Veröffentlicht: (2026)
von: Tian, Ye, et al.
Veröffentlicht: (2026)
A Framework for Longitudinal Health AI Agents
von: Lin, Georgianna, et al.
Veröffentlicht: (2026)
von: Lin, Georgianna, et al.
Veröffentlicht: (2026)
Dual‐Polarization Topological Fano Resonances Enabled by a Polarization‐Independent Topological Corner State
von: Zhuo‐Xun Peng, et al.
Veröffentlicht: (2026)
von: Zhuo‐Xun Peng, et al.
Veröffentlicht: (2026)
DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation
von: Hu, Zhenyu, et al.
Veröffentlicht: (2026)
von: Hu, Zhenyu, et al.
Veröffentlicht: (2026)
EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs
von: Liu, Shaoyu, et al.
Veröffentlicht: (2025)
von: Liu, Shaoyu, et al.
Veröffentlicht: (2025)
LiveSecBench: A Dynamic and Event-Driven Safety Benchmark for Chinese Language Model Applications
von: Li, Yudong, et al.
Veröffentlicht: (2025)
von: Li, Yudong, et al.
Veröffentlicht: (2025)
SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code
von: Li, Xinghang, et al.
Veröffentlicht: (2025)
von: Li, Xinghang, et al.
Veröffentlicht: (2025)
Mobile-Bench-v2: A More Realistic and Comprehensive Benchmark for VLM-based Mobile Agents
von: Xu, Weikai, et al.
Veröffentlicht: (2025)
von: Xu, Weikai, et al.
Veröffentlicht: (2025)
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
von: Dong, Haonan, et al.
Veröffentlicht: (2026)
von: Dong, Haonan, et al.
Veröffentlicht: (2026)
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
von: Ji, Haonian, et al.
Veröffentlicht: (2025)
von: Ji, Haonian, et al.
Veröffentlicht: (2025)
DispBench: Benchmarking Disparity Estimation to Synthetic Corruptions
von: Agnihotri, Shashank, et al.
Veröffentlicht: (2025)
von: Agnihotri, Shashank, et al.
Veröffentlicht: (2025)
UniEditBench: A Unified and Cost-Effective Benchmark for Image and Video Editing via Distilled MLLMs
von: Jiang, Lifan, et al.
Veröffentlicht: (2026)
von: Jiang, Lifan, et al.
Veröffentlicht: (2026)
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
von: Kim, Haechan, et al.
Veröffentlicht: (2026)
von: Kim, Haechan, et al.
Veröffentlicht: (2026)
Xmodel-VLM: A Simple Baseline for Multimodal Vision Language Model
von: Xu, Wanting, et al.
Veröffentlicht: (2024)
von: Xu, Wanting, et al.
Veröffentlicht: (2024)
MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents
von: Wang, Luyuan, et al.
Veröffentlicht: (2024)
von: Wang, Luyuan, et al.
Veröffentlicht: (2024)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
von: Shi, Zhichao, et al.
Veröffentlicht: (2025)
von: Shi, Zhichao, et al.
Veröffentlicht: (2025)
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
von: Nguyen, Bang, et al.
Veröffentlicht: (2026)
von: Nguyen, Bang, et al.
Veröffentlicht: (2026)
How Couples' Longitudinal Work Schedule Arrangements Shape Individual Health and Sleep at Middle Adulthood
von: Wen‐Jui Han, et al.
Veröffentlicht: (2025)
von: Wen‐Jui Han, et al.
Veröffentlicht: (2025)
NetAgentBench: A State-Centric Benchmark for Evaluating Agentic Network Configuration
von: Twabi, Ahmed, et al.
Veröffentlicht: (2026)
von: Twabi, Ahmed, et al.
Veröffentlicht: (2026)
AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents
von: Hu, Lingxiang, et al.
Veröffentlicht: (2026)
von: Hu, Lingxiang, et al.
Veröffentlicht: (2026)
SGR-Bench: Benchmarking Search Agents on State-Gated Retrieval
von: Li, Ningyuan, et al.
Veröffentlicht: (2026)
von: Li, Ningyuan, et al.
Veröffentlicht: (2026)
Health Versus Indulgence: How Goal Conflict and Risk Perception Shape Compensatory Health Beliefs in Food Bundles
von: Huayu Shi, et al.
Veröffentlicht: (2025)
von: Huayu Shi, et al.
Veröffentlicht: (2025)
MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
von: Wu, Qinzhuo, et al.
Veröffentlicht: (2026)
von: Wu, Qinzhuo, et al.
Veröffentlicht: (2026)
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis
von: Zhang, Qiaohong, et al.
Veröffentlicht: (2026)
von: Zhang, Qiaohong, et al.
Veröffentlicht: (2026)
ProBench: Benchmarking GUI Agents with Accurate Process Information
von: Yang, Leyang, et al.
Veröffentlicht: (2025)
von: Yang, Leyang, et al.
Veröffentlicht: (2025)
EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents
von: Liu, Yunqi, et al.
Veröffentlicht: (2026)
von: Liu, Yunqi, et al.
Veröffentlicht: (2026)
MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual Knowledge
von: Du, Yuntao, et al.
Veröffentlicht: (2025)
von: Du, Yuntao, et al.
Veröffentlicht: (2025)
VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment
von: Sun, Shangkun, et al.
Veröffentlicht: (2024)
von: Sun, Shangkun, et al.
Veröffentlicht: (2024)
RealBench: Benchmarking Data-Driven Numerical Weather Forecasting Under Operational Conditions and Extreme Event Challenges
von: Li, Ruize, et al.
Veröffentlicht: (2026)
von: Li, Ruize, et al.
Veröffentlicht: (2026)
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
von: Yin, Sheng, et al.
Veröffentlicht: (2024)
von: Yin, Sheng, et al.
Veröffentlicht: (2024)
SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
von: Zhang, Chenkai, et al.
Veröffentlicht: (2025)
von: Zhang, Chenkai, et al.
Veröffentlicht: (2025)
MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents
von: Zhao, Pengxiang, et al.
Veröffentlicht: (2025)
von: Zhao, Pengxiang, et al.
Veröffentlicht: (2025)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FlightBench: Benchmarking Learning-based Methods for Ego-vision-based Quadrotors Navigation
von: Yu, Shu-Ang, et al.
Veröffentlicht: (2024) -
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
von: Deng, Shihan, et al.
Veröffentlicht: (2024) -
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
von: Tu, Xinming, et al.
Veröffentlicht: (2026) -
SWE-Bench 5G: Benchmarking AI Coding Agents on Telecom Network Engineering Tasks
von: Chen, Jiao, et al.
Veröffentlicht: (2026) -
HSCO-Bench: An Agent-Driven End-to-End Hardware-Software Co-design Benchmark for Systems-on-Chip
von: Tsai, Pei-Huan, et al.
Veröffentlicht: (2026)