Knowledge-enhanced Pretraining for Vision-language Pathology Foundation Model on Cancer Diagnosis

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhou, Xiao, Sun, Luoyi, He, Dexuan, Guan, Wenbin, Wang, Ge, Wang, Ruifen, Wang, Lifeng, Yuan, Xiaojun, Sun, Xin, Zhang, Ya, Sun, Kun, Wang, Yanfeng, Xie, Weidi
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918306372386816
author Zhou, Xiao
Sun, Luoyi
He, Dexuan
Guan, Wenbin
Wang, Ge
Wang, Ruifen
Wang, Lifeng
Yuan, Xiaojun
Sun, Xin
Zhang, Ya
Sun, Kun
Wang, Yanfeng
Xie, Weidi
author_facet Zhou, Xiao
Sun, Luoyi
He, Dexuan
Guan, Wenbin
Wang, Ge
Wang, Ruifen
Wang, Lifeng
Yuan, Xiaojun
Sun, Xin
Zhang, Ya
Sun, Kun
Wang, Yanfeng
Xie, Weidi
contents Vision-language foundation models have shown great promise in computational pathology but remain primarily data-driven, lacking explicit integration of medical knowledge. We introduce KEEP (KnowledgE-Enhanced Pathology), a foundation model that systematically incorporates disease knowledge into pretraining for cancer diagnosis. KEEP leverages a comprehensive disease knowledge graph encompassing 11,454 diseases and 139,143 attributes to reorganize millions of pathology image-text pairs into 143,000 semantically structured groups aligned with disease ontology hierarchies. This knowledge-enhanced pretraining aligns visual and textual representations within hierarchical semantic spaces, enabling deeper understanding of disease relationships and morphological patterns. Across 18 public benchmarks (over 14,000 whole-slide images) and 4 institutional rare cancer datasets (926 cases), KEEP consistently outperformed existing foundation models, showing substantial gains for rare subtypes. These results establish knowledge-enhanced vision-language modeling as a powerful paradigm for advancing computational pathology.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13126
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Knowledge-enhanced Pretraining for Vision-language Pathology Foundation Model on Cancer Diagnosis
Zhou, Xiao
Sun, Luoyi
He, Dexuan
Guan, Wenbin
Wang, Ge
Wang, Ruifen
Wang, Lifeng
Yuan, Xiaojun
Sun, Xin
Zhang, Ya
Sun, Kun
Wang, Yanfeng
Xie, Weidi
Image and Video Processing
Computer Vision and Pattern Recognition
Vision-language foundation models have shown great promise in computational pathology but remain primarily data-driven, lacking explicit integration of medical knowledge. We introduce KEEP (KnowledgE-Enhanced Pathology), a foundation model that systematically incorporates disease knowledge into pretraining for cancer diagnosis. KEEP leverages a comprehensive disease knowledge graph encompassing 11,454 diseases and 139,143 attributes to reorganize millions of pathology image-text pairs into 143,000 semantically structured groups aligned with disease ontology hierarchies. This knowledge-enhanced pretraining aligns visual and textual representations within hierarchical semantic spaces, enabling deeper understanding of disease relationships and morphological patterns. Across 18 public benchmarks (over 14,000 whole-slide images) and 4 institutional rare cancer datasets (926 cases), KEEP consistently outperformed existing foundation models, showing substantial gains for rare subtypes. These results establish knowledge-enhanced vision-language modeling as a powerful paradigm for advancing computational pathology.
title Knowledge-enhanced Pretraining for Vision-language Pathology Foundation Model on Cancer Diagnosis
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.13126