Knowledge-enhanced Pretraining for Vision-language Pathology Foundation Model on Cancer Diagnosis
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866918306372386816 |
|---|---|
| author | Zhou, Xiao Sun, Luoyi He, Dexuan Guan, Wenbin Wang, Ge Wang, Ruifen Wang, Lifeng Yuan, Xiaojun Sun, Xin Zhang, Ya Sun, Kun Wang, Yanfeng Xie, Weidi |
| author_facet | Zhou, Xiao Sun, Luoyi He, Dexuan Guan, Wenbin Wang, Ge Wang, Ruifen Wang, Lifeng Yuan, Xiaojun Sun, Xin Zhang, Ya Sun, Kun Wang, Yanfeng Xie, Weidi |
| contents | Vision-language foundation models have shown great promise in computational pathology but remain primarily data-driven, lacking explicit integration of medical knowledge. We introduce KEEP (KnowledgE-Enhanced Pathology), a foundation model that systematically incorporates disease knowledge into pretraining for cancer diagnosis. KEEP leverages a comprehensive disease knowledge graph encompassing 11,454 diseases and 139,143 attributes to reorganize millions of pathology image-text pairs into 143,000 semantically structured groups aligned with disease ontology hierarchies. This knowledge-enhanced pretraining aligns visual and textual representations within hierarchical semantic spaces, enabling deeper understanding of disease relationships and morphological patterns. Across 18 public benchmarks (over 14,000 whole-slide images) and 4 institutional rare cancer datasets (926 cases), KEEP consistently outperformed existing foundation models, showing substantial gains for rare subtypes. These results establish knowledge-enhanced vision-language modeling as a powerful paradigm for advancing computational pathology. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_13126 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Knowledge-enhanced Pretraining for Vision-language Pathology Foundation Model on Cancer Diagnosis Zhou, Xiao Sun, Luoyi He, Dexuan Guan, Wenbin Wang, Ge Wang, Ruifen Wang, Lifeng Yuan, Xiaojun Sun, Xin Zhang, Ya Sun, Kun Wang, Yanfeng Xie, Weidi Image and Video Processing Computer Vision and Pattern Recognition Vision-language foundation models have shown great promise in computational pathology but remain primarily data-driven, lacking explicit integration of medical knowledge. We introduce KEEP (KnowledgE-Enhanced Pathology), a foundation model that systematically incorporates disease knowledge into pretraining for cancer diagnosis. KEEP leverages a comprehensive disease knowledge graph encompassing 11,454 diseases and 139,143 attributes to reorganize millions of pathology image-text pairs into 143,000 semantically structured groups aligned with disease ontology hierarchies. This knowledge-enhanced pretraining aligns visual and textual representations within hierarchical semantic spaces, enabling deeper understanding of disease relationships and morphological patterns. Across 18 public benchmarks (over 14,000 whole-slide images) and 4 institutional rare cancer datasets (926 cases), KEEP consistently outperformed existing foundation models, showing substantial gains for rare subtypes. These results establish knowledge-enhanced vision-language modeling as a powerful paradigm for advancing computational pathology. |
| title | Knowledge-enhanced Pretraining for Vision-language Pathology Foundation Model on Cancer Diagnosis |
| topic | Image and Video Processing Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2412.13126 |