Training Large Language Models to Predict Clinical Events

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Turtel, Benjamin, Wilczewski, Paul, Skotheim, Kris
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914561392640000
author Turtel, Benjamin
Wilczewski, Paul
Skotheim, Kris
author_facet Turtel, Benjamin
Wilczewski, Paul
Skotheim, Kris
contents Longitudinal clinical notes contain rich evidence of how patients evolve over time, but converting this signal into training supervision for clinical prediction remains challenging. We extend Foresight Learning to clinical prediction by converting time-ordered MIMIC-III notes into examples consisting of past patient context, a natural-language question about a possible future event, and a label resolved from later documentation. This process yields 6,900 prediction examples from 702 admissions across medications, procedures, organ support, microbiology, and mortality. A small LoRA adapter trained on these examples improves over the prompted base model, reducing expected calibration error from 0.1269 to 0.0398 and Brier score from 0.199 to 0.145, while slightly outperforming GPT-5 point estimates on held-out questions. The approach enables reusable clinical prediction supervision from longitudinal notes without hand-engineered structured features or endpoint-specific classifiers.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12817
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Training Large Language Models to Predict Clinical Events
Turtel, Benjamin
Wilczewski, Paul
Skotheim, Kris
Machine Learning
Artificial Intelligence
Computation and Language
Longitudinal clinical notes contain rich evidence of how patients evolve over time, but converting this signal into training supervision for clinical prediction remains challenging. We extend Foresight Learning to clinical prediction by converting time-ordered MIMIC-III notes into examples consisting of past patient context, a natural-language question about a possible future event, and a label resolved from later documentation. This process yields 6,900 prediction examples from 702 admissions across medications, procedures, organ support, microbiology, and mortality. A small LoRA adapter trained on these examples improves over the prompted base model, reducing expected calibration error from 0.1269 to 0.0398 and Brier score from 0.199 to 0.145, while slightly outperforming GPT-5 point estimates on held-out questions. The approach enables reusable clinical prediction supervision from longitudinal notes without hand-engineered structured features or endpoint-specific classifiers.
title Training Large Language Models to Predict Clinical Events
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.12817