A Multimodal Vision Foundation Model for Clinical Dermatology

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Siyuan, Yu, Zhen, Primiero, Clare, Vico-Alonso, Cristina, Wang, Zhonghua, Yang, Litao, Tschandl, Philipp, Hu, Ming, Ju, Lie, Tan, Gin, Tang, Vincent, Ng, Aik Beng, Powell, David, Bonnington, Paul, See, Simon, Magnaterra, Elisabetta, Ferguson, Peter, Nguyen, Jennifer, Guitera, Pascale, Banuls, Jose, Janda, Monika, Mar, Victoria, Kittler, Harald, Soyer, H. Peter, Ge, Zongyuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912323373891584
author Yan, Siyuan
Yu, Zhen
Primiero, Clare
Vico-Alonso, Cristina
Wang, Zhonghua
Yang, Litao
Tschandl, Philipp
Hu, Ming
Ju, Lie
Tan, Gin
Tang, Vincent
Ng, Aik Beng
Powell, David
Bonnington, Paul
See, Simon
Magnaterra, Elisabetta
Ferguson, Peter
Nguyen, Jennifer
Guitera, Pascale
Banuls, Jose
Janda, Monika
Mar, Victoria
Kittler, Harald
Soyer, H. Peter
Ge, Zongyuan
author_facet Yan, Siyuan
Yu, Zhen
Primiero, Clare
Vico-Alonso, Cristina
Wang, Zhonghua
Yang, Litao
Tschandl, Philipp
Hu, Ming
Ju, Lie
Tan, Gin
Tang, Vincent
Ng, Aik Beng
Powell, David
Bonnington, Paul
See, Simon
Magnaterra, Elisabetta
Ferguson, Peter
Nguyen, Jennifer
Guitera, Pascale
Banuls, Jose
Janda, Monika
Mar, Victoria
Kittler, Harald
Soyer, H. Peter
Ge, Zongyuan
contents Diagnosing and treating skin diseases require advanced visual skills across domains and the ability to synthesize information from multiple imaging modalities. While current deep learning models excel at specific tasks like skin cancer diagnosis from dermoscopic images, they struggle to meet the complex, multimodal requirements of clinical practice. Here, we introduce PanDerm, a multimodal dermatology foundation model pretrained through self-supervised learning on over 2 million real-world skin disease images from 11 clinical institutions across 4 imaging modalities. We evaluated PanDerm on 28 diverse benchmarks, including skin cancer screening, risk stratification, differential diagnosis of common and rare skin conditions, lesion segmentation, longitudinal monitoring, and metastasis prediction and prognosis. PanDerm achieved state-of-the-art performance across all evaluated tasks, often outperforming existing models when using only 10% of labeled data. We conducted three reader studies to assess PanDerm's potential clinical utility. PanDerm outperformed clinicians by 10.2% in early-stage melanoma detection through longitudinal analysis, improved clinicians' skin cancer diagnostic accuracy by 11% on dermoscopy images, and enhanced non-dermatologist healthcare providers' differential diagnosis by 16.5% across 128 skin conditions on clinical photographs. These results demonstrate PanDerm's potential to improve patient care across diverse clinical scenarios and serve as a model for developing multimodal foundation models in other medical specialties, potentially accelerating the integration of AI support in healthcare. The code can be found at https://github.com/SiyuanYan1/PanDerm.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15038
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Multimodal Vision Foundation Model for Clinical Dermatology
Yan, Siyuan
Yu, Zhen
Primiero, Clare
Vico-Alonso, Cristina
Wang, Zhonghua
Yang, Litao
Tschandl, Philipp
Hu, Ming
Ju, Lie
Tan, Gin
Tang, Vincent
Ng, Aik Beng
Powell, David
Bonnington, Paul
See, Simon
Magnaterra, Elisabetta
Ferguson, Peter
Nguyen, Jennifer
Guitera, Pascale
Banuls, Jose
Janda, Monika
Mar, Victoria
Kittler, Harald
Soyer, H. Peter
Ge, Zongyuan
Computer Vision and Pattern Recognition
Artificial Intelligence
Diagnosing and treating skin diseases require advanced visual skills across domains and the ability to synthesize information from multiple imaging modalities. While current deep learning models excel at specific tasks like skin cancer diagnosis from dermoscopic images, they struggle to meet the complex, multimodal requirements of clinical practice. Here, we introduce PanDerm, a multimodal dermatology foundation model pretrained through self-supervised learning on over 2 million real-world skin disease images from 11 clinical institutions across 4 imaging modalities. We evaluated PanDerm on 28 diverse benchmarks, including skin cancer screening, risk stratification, differential diagnosis of common and rare skin conditions, lesion segmentation, longitudinal monitoring, and metastasis prediction and prognosis. PanDerm achieved state-of-the-art performance across all evaluated tasks, often outperforming existing models when using only 10% of labeled data. We conducted three reader studies to assess PanDerm's potential clinical utility. PanDerm outperformed clinicians by 10.2% in early-stage melanoma detection through longitudinal analysis, improved clinicians' skin cancer diagnostic accuracy by 11% on dermoscopy images, and enhanced non-dermatologist healthcare providers' differential diagnosis by 16.5% across 128 skin conditions on clinical photographs. These results demonstrate PanDerm's potential to improve patient care across diverse clinical scenarios and serve as a model for developing multimodal foundation models in other medical specialties, potentially accelerating the integration of AI support in healthcare. The code can be found at https://github.com/SiyuanYan1/PanDerm.
title A Multimodal Vision Foundation Model for Clinical Dermatology
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2410.15038