A Scalable Framework for Evaluating Health Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Mallinar, Neil, Heydari, A. Ali, Liu, Xin, Faranesh, Anthony Z., Winslow, Brent, Hammerquist, Nova, Graef, Benjamin, Speed, Cathy, Malhotra, Mark, Patel, Shwetak, Prieto, Javier L., McDuff, Daniel, Metwally, Ahmed A. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cardiovascular-Kidney-Metabolic Health: Insights from Wearables and Blood Biomarkers
by: Esmaeilpour, Zeinab, et al.
Published: (2026)
by: Esmaeilpour, Zeinab, et al.
Published: (2026)
The Human-AI Hybrid Delphi Model: A Structured Framework for Context-Rich, Expert Consensus in Complex Domains
by: Speed, Cathy, et al.
Published: (2025)
by: Speed, Cathy, et al.
Published: (2025)
Insulin Resistance Prediction From Wearables and Routine Blood Biomarkers
by: Metwally, Ahmed A., et al.
Published: (2025)
by: Metwally, Ahmed A., et al.
Published: (2025)
The Anatomy of a Personal Health Agent
by: Heydari, A. Ali, et al.
Published: (2025)
by: Heydari, A. Ali, et al.
Published: (2025)
Transforming Wearable Data into Personal Health Insights using Large Language Model Agents
by: Merrill, Mike A., et al.
Published: (2024)
by: Merrill, Mike A., et al.
Published: (2024)
A Principle-based Framework for the Development and Evaluation of Large Language Models for Health and Wellness
by: Winslow, Brent, et al.
Published: (2025)
by: Winslow, Brent, et al.
Published: (2025)
Lifestyle-Informed Personalized Blood Biomarker Prediction via Novel Representation Learning
by: Heydari, A. Ali, et al.
Published: (2024)
by: Heydari, A. Ali, et al.
Published: (2024)
Estimating Blood Pressure with a Camera: An Exploratory Study of Ambulatory Patients with Cardiovascular Disease
by: Curran, Theodore, et al.
Published: (2025)
by: Curran, Theodore, et al.
Published: (2025)
SensorLM: Learning the Language of Wearable Sensors
by: Zhang, Yuwei, et al.
Published: (2025)
by: Zhang, Yuwei, et al.
Published: (2025)
Substance over Style: Evaluating Proactive Conversational Coaching Agents
by: Srinivas, Vidya, et al.
Published: (2025)
by: Srinivas, Vidya, et al.
Published: (2025)
Intuitive and Ubiquitous Fever Monitoring Using Smartphones and Smartwatches
by: Breda, Joseph, et al.
Published: (2021)
by: Breda, Joseph, et al.
Published: (2021)
The Potential and Perils of Generative Artificial Intelligence for Quality Improvement and Patient Safety
by: Jalilian, Laleh, et al.
Published: (2024)
by: Jalilian, Laleh, et al.
Published: (2024)
Smartphone monitoring of smiling as a behavioral proxy of well-being in everyday life
by: Poh, Ming-Zher, et al.
Published: (2025)
by: Poh, Ming-Zher, et al.
Published: (2025)
Health-LLM: Large Language Models for Health Prediction via Wearable Sensor Data
by: Kim, Yubin, et al.
Published: (2024)
by: Kim, Yubin, et al.
Published: (2024)
Towards a Personal Health Large Language Model
by: Cosentino, Justin, et al.
Published: (2024)
by: Cosentino, Justin, et al.
Published: (2024)
Adaptive Cardio Load Targets for Improving Fitness and Performance
by: Phillips, Justin, et al.
Published: (2025)
by: Phillips, Justin, et al.
Published: (2025)
Assessing the nature of large language models: A caution against anthropocentrism
by: Speed, Ann
Published: (2023)
by: Speed, Ann
Published: (2023)
RADAR: Benchmarking Language Models on Imperfect Tabular Data
by: Gu, Ken, et al.
Published: (2025)
by: Gu, Ken, et al.
Published: (2025)
New Tools are Needed for Tracking Adherence to AI Model Behavioral Use Clauses
by: McDuff, Daniel, et al.
Published: (2025)
by: McDuff, Daniel, et al.
Published: (2025)
Scaling and Load-Balancing Equi-Joins
by: Metwally, Ahmed
Published: (2022)
by: Metwally, Ahmed
Published: (2022)
HEARTS: Benchmarking LLM Reasoning on Health Time Series
by: Li, Sirui, et al.
Published: (2026)
by: Li, Sirui, et al.
Published: (2026)
Passive Heart Rate Monitoring During Smartphone Use in Everyday Life
by: Liao, Shun, et al.
Published: (2025)
by: Liao, Shun, et al.
Published: (2025)
Uneven Evolution of Cognition Across Generations of Generative AI Models
by: Galatzer-Levy, Isaac, et al.
Published: (2026)
by: Galatzer-Levy, Isaac, et al.
Published: (2026)
A Fast and Minimal System to Identify Depression Using Smartphones: Explainable Machine Learning-Based Approach
by: Ahmed, Md Sabbir, et al.
Published: (2025)
by: Ahmed, Md Sabbir, et al.
Published: (2025)
SynthWorlds: Controlled Parallel Worlds for Disentangling Reasoning and Knowledge in Language Models
by: Gu, Ken, et al.
Published: (2025)
by: Gu, Ken, et al.
Published: (2025)
Scaling Wearable Foundation Models
by: Narayanswamy, Girish, et al.
Published: (2024)
by: Narayanswamy, Girish, et al.
Published: (2024)
ConvFill: Model Collaboration for Responsive Conversational Voice Agents
by: Srinivas, Vidya, et al.
Published: (2025)
by: Srinivas, Vidya, et al.
Published: (2025)
The opportunities and risks of large language models in mental health
by: Lawrence, Hannah R., et al.
Published: (2024)
by: Lawrence, Hannah R., et al.
Published: (2024)
SympCam: Remote Optical Measurement of Sympathetic Arousal
by: Braun, Björn, et al.
Published: (2024)
by: Braun, Björn, et al.
Published: (2024)
PARSE-Ego4D: Personal Action Recommendation Suggestions for Egocentric Videos
by: Abreu, Steven, et al.
Published: (2024)
by: Abreu, Steven, et al.
Published: (2024)
Non-Contact Health Monitoring During Daily Personal Care Routines
by: Ma, Xulin, et al.
Published: (2025)
by: Ma, Xulin, et al.
Published: (2025)
Benign, Tempered, or Catastrophic: A Taxonomy of Overfitting
by: Mallinar, Neil, et al.
Published: (2022)
by: Mallinar, Neil, et al.
Published: (2022)
The chemistry of interstitial waters at DSDP Site 45-395
by: McDuff, Russell E
Published: (1984)
by: McDuff, Russell E
Published: (1984)
(Table 3) Chemistry in water at DSDP Site 45-395
by: McDuff, Russell E
Published: (1984)
by: McDuff, Russell E
Published: (1984)
(Table 2) Silicon and nitrate concentrations in water samples at DSDP Hole 45-395A
by: McDuff, Russell E
Published: (1984)
by: McDuff, Russell E
Published: (1984)
(Table 3) Interstitial water elemental composition at DSDP Leg 86 Holes
by: McDuff, Russell E
Published: (1985)
by: McDuff, Russell E
Published: (1985)
What Are the Odds? Language Models Are Capable of Probabilistic Reasoning
by: Paruchuri, Akshay, et al.
Published: (2024)
by: Paruchuri, Akshay, et al.
Published: (2024)
InvThink: Premortem Reasoning for Safer Language Models
by: Kim, Yubin, et al.
Published: (2025)
by: Kim, Yubin, et al.
Published: (2025)
A Hormone-inspired Emotion Layer for Transformer language models (HELT)
by: Reda, Eslam, et al.
Published: (2026)
by: Reda, Eslam, et al.
Published: (2026)
Eigenvectors of the De Bruijn Graph Laplacian: A Natural Basis for the Cut and Cycle Space
by: Philippakis, Anthony, et al.
Published: (2024)
by: Philippakis, Anthony, et al.
Published: (2024)
Similar Items
-
Cardiovascular-Kidney-Metabolic Health: Insights from Wearables and Blood Biomarkers
by: Esmaeilpour, Zeinab, et al.
Published: (2026) -
The Human-AI Hybrid Delphi Model: A Structured Framework for Context-Rich, Expert Consensus in Complex Domains
by: Speed, Cathy, et al.
Published: (2025) -
Insulin Resistance Prediction From Wearables and Routine Blood Biomarkers
by: Metwally, Ahmed A., et al.
Published: (2025) -
The Anatomy of a Personal Health Agent
by: Heydari, A. Ali, et al.
Published: (2025) -
Transforming Wearable Data into Personal Health Insights using Large Language Model Agents
by: Merrill, Mike A., et al.
Published: (2024)