Saved in:
Bibliographic Details
Main Authors: Unell, Alyssa, Codella, Noel C. F., Preston, Sam, Argaw, Peniel, Yim, Wen-wai, Gero, Zelalem, Wong, Cliff, Jena, Rajesh, Horvitz, Eric, Hall, Amanda K., Zhong, Ruican Rachel, Li, Jiachen, Jain, Shrey, Wei, Mu, Lungren, Matthew, Poon, Hoifung
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2509.07325
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912691088523264
author Unell, Alyssa
Codella, Noel C. F.
Preston, Sam
Argaw, Peniel
Yim, Wen-wai
Gero, Zelalem
Wong, Cliff
Jena, Rajesh
Horvitz, Eric
Hall, Amanda K.
Zhong, Ruican Rachel
Li, Jiachen
Jain, Shrey
Wei, Mu
Lungren, Matthew
Poon, Hoifung
author_facet Unell, Alyssa
Codella, Noel C. F.
Preston, Sam
Argaw, Peniel
Yim, Wen-wai
Gero, Zelalem
Wong, Cliff
Jena, Rajesh
Horvitz, Eric
Hall, Amanda K.
Zhong, Ruican Rachel
Li, Jiachen
Jain, Shrey
Wei, Mu
Lungren, Matthew
Poon, Hoifung
contents The National Comprehensive Cancer Network (NCCN) provides evidence-based guidelines for cancer treatment. Translating complex patient presentations into guideline-compliant treatment recommendations is time-intensive, requires specialized expertise, and is prone to error. Advances in large language model (LLM) capabilities promise to reduce the time required to generate treatment recommendations and improve accuracy. We present an LLM agent-based approach to automatically generate guideline-concordant treatment trajectories for patients with non-small cell lung cancer (NSCLC). Our contributions are threefold. First, we construct a novel longitudinal dataset of 121 cases of NSCLC patients that includes clinical encounters, diagnostic results, and medical histories, each expertly annotated with the corresponding NCCN guideline trajectories by board-certified oncologists. Second, we demonstrate that existing LLMs possess domain-specific knowledge that enables high-quality proxy benchmark generation for both model development and evaluation, achieving strong correlation (Spearman coefficient r=0.88, RMSE = 0.08) with expert-annotated benchmarks. Third, we develop a hybrid approach combining expensive human annotations with model consistency information to create both the agent framework that predicts the relevant guidelines for a patient, as well as a meta-classifier that verifies prediction accuracy with calibrated confidence scores for treatment recommendations (AUROC=0.800), a critical capability for communicating the accuracy of outputs, custom-tailoring tradeoffs in performance, and supporting regulatory compliance. This work establishes a framework for clinically viable LLM-based guideline adherence systems that balance accuracy, interpretability, and regulatory requirements while reducing annotation costs, providing a scalable pathway toward automated clinical decision support.
format Preprint
id arxiv_https___arxiv_org_abs_2509_07325
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CancerGUIDE: Cancer Guideline Understanding via Internal Disagreement Estimation
Unell, Alyssa
Codella, Noel C. F.
Preston, Sam
Argaw, Peniel
Yim, Wen-wai
Gero, Zelalem
Wong, Cliff
Jena, Rajesh
Horvitz, Eric
Hall, Amanda K.
Zhong, Ruican Rachel
Li, Jiachen
Jain, Shrey
Wei, Mu
Lungren, Matthew
Poon, Hoifung
Machine Learning
The National Comprehensive Cancer Network (NCCN) provides evidence-based guidelines for cancer treatment. Translating complex patient presentations into guideline-compliant treatment recommendations is time-intensive, requires specialized expertise, and is prone to error. Advances in large language model (LLM) capabilities promise to reduce the time required to generate treatment recommendations and improve accuracy. We present an LLM agent-based approach to automatically generate guideline-concordant treatment trajectories for patients with non-small cell lung cancer (NSCLC). Our contributions are threefold. First, we construct a novel longitudinal dataset of 121 cases of NSCLC patients that includes clinical encounters, diagnostic results, and medical histories, each expertly annotated with the corresponding NCCN guideline trajectories by board-certified oncologists. Second, we demonstrate that existing LLMs possess domain-specific knowledge that enables high-quality proxy benchmark generation for both model development and evaluation, achieving strong correlation (Spearman coefficient r=0.88, RMSE = 0.08) with expert-annotated benchmarks. Third, we develop a hybrid approach combining expensive human annotations with model consistency information to create both the agent framework that predicts the relevant guidelines for a patient, as well as a meta-classifier that verifies prediction accuracy with calibrated confidence scores for treatment recommendations (AUROC=0.800), a critical capability for communicating the accuracy of outputs, custom-tailoring tradeoffs in performance, and supporting regulatory compliance. This work establishes a framework for clinically viable LLM-based guideline adherence systems that balance accuracy, interpretability, and regulatory requirements while reducing annotation costs, providing a scalable pathway toward automated clinical decision support.
title CancerGUIDE: Cancer Guideline Understanding via Internal Disagreement Estimation
topic Machine Learning
url https://arxiv.org/abs/2509.07325