Health-SCORE: Towards Scalable Rubrics for Improving Health-LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911399576338432 |
|---|---|
| author | Yang, Zhichao Janghorbani, Sepehr Zhang, Dongxu Han, Jun Qian, Qian Ressler II, Andrew Lyng, Gregory D. Batra, Sanjit Singh Tillman, Robert E. |
| author_facet | Yang, Zhichao Janghorbani, Sepehr Zhang, Dongxu Han, Jun Qian, Qian Ressler II, Andrew Lyng, Gregory D. Batra, Sanjit Singh Tillman, Robert E. |
| contents | Rubrics are essential for evaluating open-ended LLM responses, especially in safety-critical domains such as healthcare. However, creating high-quality and domain-specific rubrics typically requires significant human expertise time and development cost, making rubric-based evaluation and training difficult to scale. In this work, we introduce Health-SCORE, a generalizable and scalable rubric-based training and evaluation framework that substantially reduces rubric development costs without sacrificing performance. We show that Health-SCORE provides two practical benefits beyond standalone evaluation: it can be used as a structured reward signal to guide reinforcement learning with safety-aware supervision, and it can be incorporated directly into prompts to improve response quality through in-context learning. Across open-ended healthcare tasks, Health-SCORE achieves evaluation quality comparable to human-created rubrics while significantly lowering development effort, making rubric-based evaluation and training more scalable. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_18706 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Health-SCORE: Towards Scalable Rubrics for Improving Health-LLMs Yang, Zhichao Janghorbani, Sepehr Zhang, Dongxu Han, Jun Qian, Qian Ressler II, Andrew Lyng, Gregory D. Batra, Sanjit Singh Tillman, Robert E. Artificial Intelligence Machine Learning Rubrics are essential for evaluating open-ended LLM responses, especially in safety-critical domains such as healthcare. However, creating high-quality and domain-specific rubrics typically requires significant human expertise time and development cost, making rubric-based evaluation and training difficult to scale. In this work, we introduce Health-SCORE, a generalizable and scalable rubric-based training and evaluation framework that substantially reduces rubric development costs without sacrificing performance. We show that Health-SCORE provides two practical benefits beyond standalone evaluation: it can be used as a structured reward signal to guide reinforcement learning with safety-aware supervision, and it can be incorporated directly into prompts to improve response quality through in-context learning. Across open-ended healthcare tasks, Health-SCORE achieves evaluation quality comparable to human-created rubrics while significantly lowering development effort, making rubric-based evaluation and training more scalable. |
| title | Health-SCORE: Towards Scalable Rubrics for Improving Health-LLMs |
| topic | Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2601.18706 |