Evaluating Cell AI Foundation Models in Kidney Pathology with Human-in-the-Loop Enrichment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Junlin, Lu, Siqi, Cui, Can, Deng, Ruining, Yao, Tianyuan, Tao, Zhewen, Lin, Yizhe, Lionts, Marilyn, Liu, Quan, Xiong, Juming, Wang, Yu, Zhao, Shilin, Chang, Catie, Wilkes, Mitchell, Yin, Mengmeng, Yang, Haichun, Huo, Yuankai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911254709272576
author Guo, Junlin
Lu, Siqi
Cui, Can
Deng, Ruining
Yao, Tianyuan
Tao, Zhewen
Lin, Yizhe
Lionts, Marilyn
Liu, Quan
Xiong, Juming
Wang, Yu
Zhao, Shilin
Chang, Catie
Wilkes, Mitchell
Yin, Mengmeng
Yang, Haichun
Huo, Yuankai
author_facet Guo, Junlin
Lu, Siqi
Cui, Can
Deng, Ruining
Yao, Tianyuan
Tao, Zhewen
Lin, Yizhe
Lionts, Marilyn
Liu, Quan
Xiong, Juming
Wang, Yu
Zhao, Shilin
Chang, Catie
Wilkes, Mitchell
Yin, Mengmeng
Yang, Haichun
Huo, Yuankai
contents Training AI foundation models has emerged as a promising large-scale learning approach for addressing real-world healthcare challenges, including digital pathology. While many of these models have been developed for tasks like disease diagnosis and tissue quantification using extensive and diverse training datasets, their readiness for deployment on some arguably simplest tasks, such as nuclei segmentation within a single organ (e.g., the kidney), remains uncertain. This paper seeks to answer this key question, "How good are we?", by thoroughly evaluating the performance of recent cell foundation models on a curated multi-center, multi-disease, and multi-species external testing dataset. Additionally, we tackle a more challenging question, "How can we improve?", by developing and assessing human-in-the-loop data enrichment strategies aimed at enhancing model performance while minimizing the reliance on pixel-level human annotation. To address the first question, we curated a multicenter, multidisease, and multispecies dataset consisting of 2,542 kidney whole slide images (WSIs). Three state-of-the-art (SOTA) cell foundation models-Cellpose, StarDist, and CellViT-were selected for evaluation. To tackle the second question, we explored data enrichment algorithms by distilling predictions from the different foundation models with a human-in-the-loop framework, aiming to further enhance foundation model performance with minimal human efforts. Our experimental results showed that all three foundation models improved over their baselines with model fine-tuning with enriched data. Interestingly, the baseline model with the highest F1 score does not yield the best segmentation outcomes after fine-tuning. This study establishes a benchmark for the development and deployment of cell vision foundation models tailored for real-world data applications.
format Preprint
id arxiv_https___arxiv_org_abs_2411_00078
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating Cell AI Foundation Models in Kidney Pathology with Human-in-the-Loop Enrichment
Guo, Junlin
Lu, Siqi
Cui, Can
Deng, Ruining
Yao, Tianyuan
Tao, Zhewen
Lin, Yizhe
Lionts, Marilyn
Liu, Quan
Xiong, Juming
Wang, Yu
Zhao, Shilin
Chang, Catie
Wilkes, Mitchell
Yin, Mengmeng
Yang, Haichun
Huo, Yuankai
Computer Vision and Pattern Recognition
Artificial Intelligence
Image and Video Processing
Training AI foundation models has emerged as a promising large-scale learning approach for addressing real-world healthcare challenges, including digital pathology. While many of these models have been developed for tasks like disease diagnosis and tissue quantification using extensive and diverse training datasets, their readiness for deployment on some arguably simplest tasks, such as nuclei segmentation within a single organ (e.g., the kidney), remains uncertain. This paper seeks to answer this key question, "How good are we?", by thoroughly evaluating the performance of recent cell foundation models on a curated multi-center, multi-disease, and multi-species external testing dataset. Additionally, we tackle a more challenging question, "How can we improve?", by developing and assessing human-in-the-loop data enrichment strategies aimed at enhancing model performance while minimizing the reliance on pixel-level human annotation. To address the first question, we curated a multicenter, multidisease, and multispecies dataset consisting of 2,542 kidney whole slide images (WSIs). Three state-of-the-art (SOTA) cell foundation models-Cellpose, StarDist, and CellViT-were selected for evaluation. To tackle the second question, we explored data enrichment algorithms by distilling predictions from the different foundation models with a human-in-the-loop framework, aiming to further enhance foundation model performance with minimal human efforts. Our experimental results showed that all three foundation models improved over their baselines with model fine-tuning with enriched data. Interestingly, the baseline model with the highest F1 score does not yield the best segmentation outcomes after fine-tuning. This study establishes a benchmark for the development and deployment of cell vision foundation models tailored for real-world data applications.
title Evaluating Cell AI Foundation Models in Kidney Pathology with Human-in-the-Loop Enrichment
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Image and Video Processing
url https://arxiv.org/abs/2411.00078