LifeAgentBench: A Multi-dimensional Benchmark and Agent for Personal Health Assistants in Digital Health
Fuente:
arXiv
Guardado en:
| Autores principales: | Tian, Ye, Wang, Zihao, Gungor, Onat, Fan, Xiaoran, Rosing, Tajana |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FoodCHA: Multi-Modal LLM Agent for Fine-Grained Food Analysis
por: Lee, Woojin, et al.
Publicado: (2026)
por: Lee, Woojin, et al.
Publicado: (2026)
DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs
por: Tian, Ye, et al.
Publicado: (2025)
por: Tian, Ye, et al.
Publicado: (2025)
KLDrive: Fine-Grained 3D Scene Reasoning for Autonomous Driving based on Knowledge Graph
por: Tian, Ye, et al.
Publicado: (2026)
por: Tian, Ye, et al.
Publicado: (2026)
E-QUARTIC: Energy Efficient Edge Ensemble of Convolutional Neural Networks for Resource-Optimized Learning
por: Zhang, Le, et al.
Publicado: (2024)
por: Zhang, Le, et al.
Publicado: (2024)
ACORN-IDS: Adaptive Continual Novelty Detection for Intrusion Detection Systems
por: Fuhrman, Sean, et al.
Publicado: (2026)
por: Fuhrman, Sean, et al.
Publicado: (2026)
CND-IDS: Continual Novelty Detection for Intrusion Detection Systems
por: Fuhrman, Sean, et al.
Publicado: (2025)
por: Fuhrman, Sean, et al.
Publicado: (2025)
CAN-QA: A Question-Answering Benchmark for Reasoning over In-Vehicle CAN Traffic
por: Chen, Jing, et al.
Publicado: (2026)
por: Chen, Jing, et al.
Publicado: (2026)
CyberMaskQA: A Privacy-Aware Benchmark for Evaluating Large Language Models in Cybersecurity Question Answering
por: Gaddi, Matilda, et al.
Publicado: (2026)
por: Gaddi, Matilda, et al.
Publicado: (2026)
Offload Rethinking by Cloud Assistance for Efficient Environmental Sound Recognition on LPWANs
por: Zhang, Le, et al.
Publicado: (2025)
por: Zhang, Le, et al.
Publicado: (2025)
$π$-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
por: Zhang, Haoran, et al.
Publicado: (2026)
por: Zhang, Haoran, et al.
Publicado: (2026)
AQUA-LLM: Evaluating Accuracy, Quantization, and Adversarial Robustness Trade-offs in LLMs for Cybersecurity Question Answering
por: Gungor, Onat, et al.
Publicado: (2025)
por: Gungor, Onat, et al.
Publicado: (2025)
LIGHT-HIDS: A Lightweight and Effective Machine Learning-Based Framework for Robust Host Intrusion Detection
por: Gungor, Onat, et al.
Publicado: (2025)
por: Gungor, Onat, et al.
Publicado: (2025)
CITADEL: Continual Anomaly Detection for Enhanced Learning in IoT Intrusion Detection
por: Li, Elvin, et al.
Publicado: (2025)
por: Li, Elvin, et al.
Publicado: (2025)
EAGER: Edge-Aligned LLM Defense for Robust, Efficient, and Accurate Cybersecurity Question Answering
por: Gungor, Onat, et al.
Publicado: (2025)
por: Gungor, Onat, et al.
Publicado: (2025)
SAGE: Sample-Aware Guarding Engine for Robust Intrusion Detection Against Adversarial Attacks
por: Chen, Jing, et al.
Publicado: (2025)
por: Chen, Jing, et al.
Publicado: (2025)
SAFE: Self-Supervised Anomaly Detection Framework for Intrusion Detection
por: Li, Elvin, et al.
Publicado: (2025)
por: Li, Elvin, et al.
Publicado: (2025)
ESL-Bench: An Event-Driven Synthetic Longitudinal Benchmark for Health Agents
por: Li, Chao, et al.
Publicado: (2026)
por: Li, Chao, et al.
Publicado: (2026)
PSPA-Bench: A Personalized Benchmark for Smartphone GUI Agent
por: Nie, Hongyi, et al.
Publicado: (2026)
por: Nie, Hongyi, et al.
Publicado: (2026)
MicroHD: An Accuracy-Driven Optimization of Hyperdimensional Computing Algorithms for TinyML systems
por: Ponzina, Flavio, et al.
Publicado: (2024)
por: Ponzina, Flavio, et al.
Publicado: (2024)
INTARG: Informed Real-Time Adversarial Attack Generation for Time-Series Regression
por: Tokgoz, Gamze Kirman, et al.
Publicado: (2026)
por: Tokgoz, Gamze Kirman, et al.
Publicado: (2026)
TS-OOD: Evaluating Time-Series Out-of-Distribution Detection and Prospective Directions for Progress
por: Gungor, Onat, et al.
Publicado: (2025)
por: Gungor, Onat, et al.
Publicado: (2025)
ReLATE+: Unified Framework for Adversarial Attack Detection, Classification, and Resilient Model Selection in Time-Series Classification
por: Kocal, Cagla Ipek, et al.
Publicado: (2025)
por: Kocal, Cagla Ipek, et al.
Publicado: (2025)
From Assistant to Double Agent: Formalizing and Benchmarking Attacks on OpenClaw for Personalized Local AI Agent
por: Wang, Yuhang, et al.
Publicado: (2026)
por: Wang, Yuhang, et al.
Publicado: (2026)
SensorQA: A Question Answering Benchmark for Daily-Life Monitoring
por: Reichman, Benjamin, et al.
Publicado: (2025)
por: Reichman, Benjamin, et al.
Publicado: (2025)
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
por: Dong, Haonan, et al.
Publicado: (2026)
por: Dong, Haonan, et al.
Publicado: (2026)
AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems
por: Shang, Yu, et al.
Publicado: (2025)
por: Shang, Yu, et al.
Publicado: (2025)
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs
por: Hu, Lanxiang, et al.
Publicado: (2024)
por: Hu, Lanxiang, et al.
Publicado: (2024)
LifeBench: A Benchmark for Long-Horizon Multi-Source Memory
por: Cheng, Zihao, et al.
Publicado: (2026)
por: Cheng, Zihao, et al.
Publicado: (2026)
DYNAMITE: Dynamic Defense Selection for Enhancing Machine Learning-based Intrusion Detection Against Adversarial Attacks
por: Chen, Jing, et al.
Publicado: (2025)
por: Chen, Jing, et al.
Publicado: (2025)
Small Agent Group is the Future of Digital Health
por: Meng, Yuqiao, et al.
Publicado: (2026)
por: Meng, Yuqiao, et al.
Publicado: (2026)
OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents
por: Hu, Yulin, et al.
Publicado: (2026)
por: Hu, Yulin, et al.
Publicado: (2026)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
por: Long, Xiang, et al.
Publicado: (2026)
por: Long, Xiang, et al.
Publicado: (2026)
Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World
por: Lin, Yusong, et al.
Publicado: (2026)
por: Lin, Yusong, et al.
Publicado: (2026)
MASLegalBench: Benchmarking Multi-Agent Systems in Deductive Legal Reasoning
por: Jing, Huihao, et al.
Publicado: (2025)
por: Jing, Huihao, et al.
Publicado: (2025)
Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench
por: Hu, Lanxiang, et al.
Publicado: (2025)
por: Hu, Lanxiang, et al.
Publicado: (2025)
M^3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
por: Zhou, Yang, et al.
Publicado: (2025)
por: Zhou, Yang, et al.
Publicado: (2025)
BenchAgents: Multi-Agent Systems for Structured Benchmark Creation
por: Butt, Natasha, et al.
Publicado: (2024)
por: Butt, Natasha, et al.
Publicado: (2024)
The Anatomy of a Personal Health Agent
por: Heydari, A. Ali, et al.
Publicado: (2025)
por: Heydari, A. Ali, et al.
Publicado: (2025)
lmgame-Bench: How Good are LLMs at Playing Games?
por: Hu, Lanxiang, et al.
Publicado: (2025)
por: Hu, Lanxiang, et al.
Publicado: (2025)
A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models
por: Kong, Jason, et al.
Publicado: (2026)
por: Kong, Jason, et al.
Publicado: (2026)
Ejemplares similares
-
FoodCHA: Multi-Modal LLM Agent for Fine-Grained Food Analysis
por: Lee, Woojin, et al.
Publicado: (2026) -
DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs
por: Tian, Ye, et al.
Publicado: (2025) -
KLDrive: Fine-Grained 3D Scene Reasoning for Autonomous Driving based on Knowledge Graph
por: Tian, Ye, et al.
Publicado: (2026) -
E-QUARTIC: Energy Efficient Edge Ensemble of Convolutional Neural Networks for Resource-Optimized Learning
por: Zhang, Le, et al.
Publicado: (2024) -
ACORN-IDS: Adaptive Continual Novelty Detection for Intrusion Detection Systems
por: Fuhrman, Sean, et al.
Publicado: (2026)