Abnormality-Driven Representation Learning for Radiology Imaging

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ligero, Marta, Lenz, Tim, Wölflein, Georg, Nahhas, Omar S. M. El, Truhn, Daniel, Kather, Jakob Nikolas
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909404122578944
author Ligero, Marta
Lenz, Tim
Wölflein, Georg
Nahhas, Omar S. M. El
Truhn, Daniel
Kather, Jakob Nikolas
author_facet Ligero, Marta
Lenz, Tim
Wölflein, Georg
Nahhas, Omar S. M. El
Truhn, Daniel
Kather, Jakob Nikolas
contents To date, the most common approach for radiology deep learning pipelines is the use of end-to-end 3D networks based on models pre-trained on other tasks, followed by fine-tuning on the task at hand. In contrast, adjacent medical fields such as pathology, which focus on 2D images, have effectively adopted task-agnostic foundational models based on self-supervised learning (SSL), combined with weakly-supervised deep learning (DL). However, the field of radiology still lacks task-agnostic representation models due to the computational and data demands of 3D imaging and the anatomical complexity inherent to radiology scans. To address this gap, we propose CLEAR, a framework for radiology images that uses extracted embeddings from 2D slices along with attention-based aggregation for efficiently predicting clinical endpoints. As part of this framework, we introduce lesion-enhanced contrastive learning (LeCL), a novel approach to obtain visual representations driven by abnormalities in 2D axial slices across different locations of the CT scans. Specifically, we trained single-domain contrastive learning approaches using three different architectures: Vision Transformers, Vision State Space Models and Gated Convolutional Neural Networks. We evaluate our approach across three clinical tasks: tumor lesion location, lung disease detection, and patient staging, benchmarking against four state-of-the-art foundation models, including BiomedCLIP. Our findings demonstrate that CLEAR using representations learned through LeCL, outperforms existing foundation models, while being substantially more compute- and data-efficient.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16803
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Abnormality-Driven Representation Learning for Radiology Imaging
Ligero, Marta
Lenz, Tim
Wölflein, Georg
Nahhas, Omar S. M. El
Truhn, Daniel
Kather, Jakob Nikolas
Computer Vision and Pattern Recognition
To date, the most common approach for radiology deep learning pipelines is the use of end-to-end 3D networks based on models pre-trained on other tasks, followed by fine-tuning on the task at hand. In contrast, adjacent medical fields such as pathology, which focus on 2D images, have effectively adopted task-agnostic foundational models based on self-supervised learning (SSL), combined with weakly-supervised deep learning (DL). However, the field of radiology still lacks task-agnostic representation models due to the computational and data demands of 3D imaging and the anatomical complexity inherent to radiology scans. To address this gap, we propose CLEAR, a framework for radiology images that uses extracted embeddings from 2D slices along with attention-based aggregation for efficiently predicting clinical endpoints. As part of this framework, we introduce lesion-enhanced contrastive learning (LeCL), a novel approach to obtain visual representations driven by abnormalities in 2D axial slices across different locations of the CT scans. Specifically, we trained single-domain contrastive learning approaches using three different architectures: Vision Transformers, Vision State Space Models and Gated Convolutional Neural Networks. We evaluate our approach across three clinical tasks: tumor lesion location, lung disease detection, and patient staging, benchmarking against four state-of-the-art foundation models, including BiomedCLIP. Our findings demonstrate that CLEAR using representations learned through LeCL, outperforms existing foundation models, while being substantially more compute- and data-efficient.
title Abnormality-Driven Representation Learning for Radiology Imaging
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.16803