Understanding differences in applying DETR to natural and medical images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Yanqi, Shen, Yiqiu, Fernandez-Granda, Carlos, Heacock, Laura, Geras, Krzysztof J.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915314355142656
author Xu, Yanqi
Shen, Yiqiu
Fernandez-Granda, Carlos
Heacock, Laura
Geras, Krzysztof J.
author_facet Xu, Yanqi
Shen, Yiqiu
Fernandez-Granda, Carlos
Heacock, Laura
Geras, Krzysztof J.
contents Transformer-based detectors have shown success in computer vision tasks with natural images. These models, exemplified by the Deformable DETR, are optimized through complex engineering strategies tailored to the typical characteristics of natural scenes. However, medical imaging data presents unique challenges such as extremely large image sizes, fewer and smaller regions of interest, and object classes which can be differentiated only through subtle differences. This study evaluates the applicability of these transformer-based design choices when applied to a screening mammography dataset that represents these distinct medical imaging data characteristics. Our analysis reveals that common design choices from the natural image domain, such as complex encoder architectures, multi-scale feature fusion, query initialization, and iterative bounding box refinement, do not improve and sometimes even impair object detection performance in medical imaging. In contrast, simpler and shallower architectures often achieve equal or superior results. This finding suggests that the adaptation of transformer models for medical imaging data requires a reevaluation of standard practices, potentially leading to more efficient and specialized frameworks for medical diagnosis.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17677
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Understanding differences in applying DETR to natural and medical images
Xu, Yanqi
Shen, Yiqiu
Fernandez-Granda, Carlos
Heacock, Laura
Geras, Krzysztof J.
Computer Vision and Pattern Recognition
Transformer-based detectors have shown success in computer vision tasks with natural images. These models, exemplified by the Deformable DETR, are optimized through complex engineering strategies tailored to the typical characteristics of natural scenes. However, medical imaging data presents unique challenges such as extremely large image sizes, fewer and smaller regions of interest, and object classes which can be differentiated only through subtle differences. This study evaluates the applicability of these transformer-based design choices when applied to a screening mammography dataset that represents these distinct medical imaging data characteristics. Our analysis reveals that common design choices from the natural image domain, such as complex encoder architectures, multi-scale feature fusion, query initialization, and iterative bounding box refinement, do not improve and sometimes even impair object detection performance in medical imaging. In contrast, simpler and shallower architectures often achieve equal or superior results. This finding suggests that the adaptation of transformer models for medical imaging data requires a reevaluation of standard practices, potentially leading to more efficient and specialized frameworks for medical diagnosis.
title Understanding differences in applying DETR to natural and medical images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.17677