SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understanding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Ying, Wang, Guoan, Ji, Yuanfeng, Li, Yanjun, Ye, Jin, Li, Tianbin, Hu, Ming, Yu, Rongshan, Qiao, Yu, He, Junjun
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908275336806400
author Chen, Ying
Wang, Guoan
Ji, Yuanfeng
Li, Yanjun
Ye, Jin
Li, Tianbin
Hu, Ming
Yu, Rongshan
Qiao, Yu
He, Junjun
author_facet Chen, Ying
Wang, Guoan
Ji, Yuanfeng
Li, Yanjun
Ye, Jin
Li, Tianbin
Hu, Ming
Yu, Rongshan
Qiao, Yu
He, Junjun
contents Despite the progress made by multimodal large language models (MLLMs) in computational pathology, they remain limited by a predominant focus on patch-level analysis, missing essential contextual information at the whole-slide level. The lack of large-scale instruction datasets and the gigapixel scale of whole slide images (WSIs) pose significant developmental challenges. In this paper, we present SlideChat, the first vision-language assistant capable of understanding gigapixel whole-slide images, exhibiting excellent multimodal conversational capability and response complex instruction across diverse pathology scenarios. To support its development, we created SlideInstruction, the largest instruction-following dataset for WSIs consisting of 4.2K WSI captions and 176K VQA pairs with multiple categories. Furthermore, we propose SlideBench, a multimodal benchmark that incorporates captioning and VQA tasks to assess SlideChat's capabilities in varied clinical settings such as microscopy, diagnosis. Compared to both general and specialized MLLMs, SlideChat exhibits exceptional capabilities achieving state-of-the-art performance on 18 of 22 tasks. For example, it achieved an overall accuracy of 81.17% on SlideBench-VQA (TCGA), and 54.15% on SlideBench-VQA (BCNB). Our code, data, and model is publicly accessible at https://uni-medical.github.io/SlideChat.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11761
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understanding
Chen, Ying
Wang, Guoan
Ji, Yuanfeng
Li, Yanjun
Ye, Jin
Li, Tianbin
Hu, Ming
Yu, Rongshan
Qiao, Yu
He, Junjun
Computer Vision and Pattern Recognition
Artificial Intelligence
Despite the progress made by multimodal large language models (MLLMs) in computational pathology, they remain limited by a predominant focus on patch-level analysis, missing essential contextual information at the whole-slide level. The lack of large-scale instruction datasets and the gigapixel scale of whole slide images (WSIs) pose significant developmental challenges. In this paper, we present SlideChat, the first vision-language assistant capable of understanding gigapixel whole-slide images, exhibiting excellent multimodal conversational capability and response complex instruction across diverse pathology scenarios. To support its development, we created SlideInstruction, the largest instruction-following dataset for WSIs consisting of 4.2K WSI captions and 176K VQA pairs with multiple categories. Furthermore, we propose SlideBench, a multimodal benchmark that incorporates captioning and VQA tasks to assess SlideChat's capabilities in varied clinical settings such as microscopy, diagnosis. Compared to both general and specialized MLLMs, SlideChat exhibits exceptional capabilities achieving state-of-the-art performance on 18 of 22 tasks. For example, it achieved an overall accuracy of 81.17% on SlideBench-VQA (TCGA), and 54.15% on SlideBench-VQA (BCNB). Our code, data, and model is publicly accessible at https://uni-medical.github.io/SlideChat.github.io.
title SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2410.11761