PD-L1 Classification of Weakly-Labeled Whole Slide Images of Breast Cancer

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cignoni, Giacomo, Scatena, Cristian, Frascarelli, Chiara, Fusco, Nicola, Naccarato, Antonio Giuseppe, Fanelli, Giuseppe Nicoló, Sîrbu, Alina
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909171225460736
author Cignoni, Giacomo
Scatena, Cristian
Frascarelli, Chiara
Fusco, Nicola
Naccarato, Antonio Giuseppe
Fanelli, Giuseppe Nicoló
Sîrbu, Alina
author_facet Cignoni, Giacomo
Scatena, Cristian
Frascarelli, Chiara
Fusco, Nicola
Naccarato, Antonio Giuseppe
Fanelli, Giuseppe Nicoló
Sîrbu, Alina
contents Specific and effective breast cancer therapy relies on the accurate quantification of PD-L1 positivity in tumors, which appears in the form of brown stainings in high resolution whole slide images (WSIs). However, the retrieval and extensive labeling of PD-L1 stained WSIs is a time-consuming and challenging task for pathologists, resulting in low reproducibility, especially for borderline images. This study aims to develop and compare models able to classify PD-L1 positivity of breast cancer samples based on WSI analysis, relying only on WSI-level labels. The task consists of two phases: identifying regions of interest (ROI) and classifying tumors as PD-L1 positive or negative. For the latter, two model categories were developed, with different feature extraction methodologies. The first encodes images based on the colour distance from a base color. The second uses a convolutional autoencoder to obtain embeddings of WSI tiles, and aggregates them into a WSI-level embedding. For both model types, features are fed into downstream ML classifiers. Two datasets from different clinical centers were used in two different training configurations: (1) training on one dataset and testing on the other; (2) combining the datasets. We also tested the performance with or without human preprocessing to remove brown artefacts Colour distance based models achieve the best performances on testing configuration (1) with artefact removal, while autoencoder-based models are superior in the remaining cases, which are prone to greater data variability.
format Preprint
id arxiv_https___arxiv_org_abs_2404_10175
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PD-L1 Classification of Weakly-Labeled Whole Slide Images of Breast Cancer
Cignoni, Giacomo
Scatena, Cristian
Frascarelli, Chiara
Fusco, Nicola
Naccarato, Antonio Giuseppe
Fanelli, Giuseppe Nicoló
Sîrbu, Alina
Computational Engineering, Finance, and Science
Computer Vision and Pattern Recognition
J.3; I.4.9
Specific and effective breast cancer therapy relies on the accurate quantification of PD-L1 positivity in tumors, which appears in the form of brown stainings in high resolution whole slide images (WSIs). However, the retrieval and extensive labeling of PD-L1 stained WSIs is a time-consuming and challenging task for pathologists, resulting in low reproducibility, especially for borderline images. This study aims to develop and compare models able to classify PD-L1 positivity of breast cancer samples based on WSI analysis, relying only on WSI-level labels. The task consists of two phases: identifying regions of interest (ROI) and classifying tumors as PD-L1 positive or negative. For the latter, two model categories were developed, with different feature extraction methodologies. The first encodes images based on the colour distance from a base color. The second uses a convolutional autoencoder to obtain embeddings of WSI tiles, and aggregates them into a WSI-level embedding. For both model types, features are fed into downstream ML classifiers. Two datasets from different clinical centers were used in two different training configurations: (1) training on one dataset and testing on the other; (2) combining the datasets. We also tested the performance with or without human preprocessing to remove brown artefacts Colour distance based models achieve the best performances on testing configuration (1) with artefact removal, while autoencoder-based models are superior in the remaining cases, which are prone to greater data variability.
title PD-L1 Classification of Weakly-Labeled Whole Slide Images of Breast Cancer
topic Computational Engineering, Finance, and Science
Computer Vision and Pattern Recognition
J.3; I.4.9
url https://arxiv.org/abs/2404.10175