KEPIL: Knowledge-Enhanced Prompt-Image Learning for Prompt-Robust Disease Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Haozhe, Shu, Shelley Zixin, Zhou, Ziyu, Berke, Robert, Reyes, Mauricio
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911667406766080
author Luo, Haozhe
Shu, Shelley Zixin
Zhou, Ziyu
Berke, Robert
Reyes, Mauricio
author_facet Luo, Haozhe
Shu, Shelley Zixin
Zhou, Ziyu
Berke, Robert
Reyes, Mauricio
contents Vision--language models (VLMs) show promise for clinical decision support in radiology because they enable joint reasoning over radiological images and clinical text, thereby leveraging complementary clinical information. However, radiological findings are long-tailed in practice, leaving some conditions underrepresented and making zero-shot inference essential. Yet current CLIP-style medical VLMs are sensitive to prompt variations and often lack trustworthy external knowledge at inference time, which hinders reliable clinical deployment. We present \textit{KEPIL}, a prompt-robust framework that integrates curated medical knowledge to stabilize zero-shot generalization. KEPIL comprises: (i) \emph{dynamic prompt enrichment} using ontologies with LLM assistance, (ii) a \emph{semantic-aware contrastive loss} aligning embeddings of equivalent prompt variants via a dual-embedding objective, and (iii) \emph{entity-centric report standardization} to yield ontology-aligned representations. Across seven benchmarks, KEPIL achieves state-of-the-art zero-shot inference performance; under prompt-variation tests, it improves AUC by \(6.37\%\) on \textit{CheXpert} and by \(4.11\%\) on average. These results suggest that structured knowledge and robust prompt design are key to clinically reliable radiology-facing VLMs. Code will be released at https://github.com/Roypic/KEPIL.
format Preprint
id arxiv_https___arxiv_org_abs_2605_09132
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle KEPIL: Knowledge-Enhanced Prompt-Image Learning for Prompt-Robust Disease Detection
Luo, Haozhe
Shu, Shelley Zixin
Zhou, Ziyu
Berke, Robert
Reyes, Mauricio
Computer Vision and Pattern Recognition
Vision--language models (VLMs) show promise for clinical decision support in radiology because they enable joint reasoning over radiological images and clinical text, thereby leveraging complementary clinical information. However, radiological findings are long-tailed in practice, leaving some conditions underrepresented and making zero-shot inference essential. Yet current CLIP-style medical VLMs are sensitive to prompt variations and often lack trustworthy external knowledge at inference time, which hinders reliable clinical deployment. We present \textit{KEPIL}, a prompt-robust framework that integrates curated medical knowledge to stabilize zero-shot generalization. KEPIL comprises: (i) \emph{dynamic prompt enrichment} using ontologies with LLM assistance, (ii) a \emph{semantic-aware contrastive loss} aligning embeddings of equivalent prompt variants via a dual-embedding objective, and (iii) \emph{entity-centric report standardization} to yield ontology-aligned representations. Across seven benchmarks, KEPIL achieves state-of-the-art zero-shot inference performance; under prompt-variation tests, it improves AUC by \(6.37\%\) on \textit{CheXpert} and by \(4.11\%\) on average. These results suggest that structured knowledge and robust prompt design are key to clinically reliable radiology-facing VLMs. Code will be released at https://github.com/Roypic/KEPIL.
title KEPIL: Knowledge-Enhanced Prompt-Image Learning for Prompt-Robust Disease Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.09132