Automatic Construction of Clinical Scoring Systems with LLM Agents

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Estévez, Silas Ruhrberg, Chiu, Christopher, van der Schaar, Mihaela
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914589284761600
author Estévez, Silas Ruhrberg
Chiu, Christopher
van der Schaar, Mihaela
author_facet Estévez, Silas Ruhrberg
Chiu, Christopher
van der Schaar, Mihaela
contents Modern clinical practice relies on evidence-based guidelines implemented as compact scoring systems composed of a small number of interpretable decision rules. While machine-learning models achieve strong performance, many fail to translate into routine clinical use due to misalignment with workflow constraints such as memorability, auditability, and bedside execution. We argue that this gap arises not from insufficient predictive power, but from optimizing over model classes that are incompatible with guideline deployment. Deployable guidelines often take the form of unit-weighted clinical checklists, formed by thresholding the sum of binary rules, but learning such scores requires searching an exponentially large discrete space of possible rule sets. We introduce AgentScore, which performs semantically guided optimization in this space by using LLMs to propose candidate rules and a deterministic, data-grounded verification-and-selection loop to enforce statistical validity and deployability constraints. Across eight clinical prediction tasks, AgentScore outperforms existing score-generation methods and achieves AUROC comparable to more flexible interpretable models despite operating under stronger structural constraints. On two additional externally validated tasks, AgentScore achieves higher discrimination than established guideline-based scores.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22324
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Automatic Construction of Clinical Scoring Systems with LLM Agents
Estévez, Silas Ruhrberg
Chiu, Christopher
van der Schaar, Mihaela
Machine Learning
Multiagent Systems
Modern clinical practice relies on evidence-based guidelines implemented as compact scoring systems composed of a small number of interpretable decision rules. While machine-learning models achieve strong performance, many fail to translate into routine clinical use due to misalignment with workflow constraints such as memorability, auditability, and bedside execution. We argue that this gap arises not from insufficient predictive power, but from optimizing over model classes that are incompatible with guideline deployment. Deployable guidelines often take the form of unit-weighted clinical checklists, formed by thresholding the sum of binary rules, but learning such scores requires searching an exponentially large discrete space of possible rule sets. We introduce AgentScore, which performs semantically guided optimization in this space by using LLMs to propose candidate rules and a deterministic, data-grounded verification-and-selection loop to enforce statistical validity and deployability constraints. Across eight clinical prediction tasks, AgentScore outperforms existing score-generation methods and achieves AUROC comparable to more flexible interpretable models despite operating under stronger structural constraints. On two additional externally validated tasks, AgentScore achieves higher discrimination than established guideline-based scores.
title Automatic Construction of Clinical Scoring Systems with LLM Agents
topic Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2601.22324