Health-SCORE: Towards Scalable Rubrics for Improving Health-LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Zhichao, Janghorbani, Sepehr, Zhang, Dongxu, Han, Jun, Qian, Qian, Ressler II, Andrew, Lyng, Gregory D., Batra, Sanjit Singh, Tillman, Robert E.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911399576338432
author Yang, Zhichao
Janghorbani, Sepehr
Zhang, Dongxu
Han, Jun
Qian, Qian
Ressler II, Andrew
Lyng, Gregory D.
Batra, Sanjit Singh
Tillman, Robert E.
author_facet Yang, Zhichao
Janghorbani, Sepehr
Zhang, Dongxu
Han, Jun
Qian, Qian
Ressler II, Andrew
Lyng, Gregory D.
Batra, Sanjit Singh
Tillman, Robert E.
contents Rubrics are essential for evaluating open-ended LLM responses, especially in safety-critical domains such as healthcare. However, creating high-quality and domain-specific rubrics typically requires significant human expertise time and development cost, making rubric-based evaluation and training difficult to scale. In this work, we introduce Health-SCORE, a generalizable and scalable rubric-based training and evaluation framework that substantially reduces rubric development costs without sacrificing performance. We show that Health-SCORE provides two practical benefits beyond standalone evaluation: it can be used as a structured reward signal to guide reinforcement learning with safety-aware supervision, and it can be incorporated directly into prompts to improve response quality through in-context learning. Across open-ended healthcare tasks, Health-SCORE achieves evaluation quality comparable to human-created rubrics while significantly lowering development effort, making rubric-based evaluation and training more scalable.
format Preprint
id arxiv_https___arxiv_org_abs_2601_18706
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Health-SCORE: Towards Scalable Rubrics for Improving Health-LLMs
Yang, Zhichao
Janghorbani, Sepehr
Zhang, Dongxu
Han, Jun
Qian, Qian
Ressler II, Andrew
Lyng, Gregory D.
Batra, Sanjit Singh
Tillman, Robert E.
Artificial Intelligence
Machine Learning
Rubrics are essential for evaluating open-ended LLM responses, especially in safety-critical domains such as healthcare. However, creating high-quality and domain-specific rubrics typically requires significant human expertise time and development cost, making rubric-based evaluation and training difficult to scale. In this work, we introduce Health-SCORE, a generalizable and scalable rubric-based training and evaluation framework that substantially reduces rubric development costs without sacrificing performance. We show that Health-SCORE provides two practical benefits beyond standalone evaluation: it can be used as a structured reward signal to guide reinforcement learning with safety-aware supervision, and it can be incorporated directly into prompts to improve response quality through in-context learning. Across open-ended healthcare tasks, Health-SCORE achieves evaluation quality comparable to human-created rubrics while significantly lowering development effort, making rubric-based evaluation and training more scalable.
title Health-SCORE: Towards Scalable Rubrics for Improving Health-LLMs
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2601.18706