DuoFormer: Leveraging Hierarchical Visual Representations by Local and Global Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Xiaoya, Zhang, Bodong, Knudsen, Beatrice S., Tasdizen, Tolga
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911961793429504
author Tang, Xiaoya
Zhang, Bodong
Knudsen, Beatrice S.
Tasdizen, Tolga
author_facet Tang, Xiaoya
Zhang, Bodong
Knudsen, Beatrice S.
Tasdizen, Tolga
contents We here propose a novel hierarchical transformer model that adeptly integrates the feature extraction capabilities of Convolutional Neural Networks (CNNs) with the advanced representational potential of Vision Transformers (ViTs). Addressing the lack of inductive biases and dependence on extensive training datasets in ViTs, our model employs a CNN backbone to generate hierarchical visual representations. These representations are then adapted for transformer input through an innovative patch tokenization. We also introduce a 'scale attention' mechanism that captures cross-scale dependencies, complementing patch attention to enhance spatial understanding and preserve global perception. Our approach significantly outperforms baseline models on small and medium-sized medical datasets, demonstrating its efficiency and generalizability. The components are designed as plug-and-play for different CNN architectures and can be adapted for multiple applications. The code is available at https://github.com/xiaoyatang/DuoFormer.git.
format Preprint
id arxiv_https___arxiv_org_abs_2407_13920
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DuoFormer: Leveraging Hierarchical Visual Representations by Local and Global Attention
Tang, Xiaoya
Zhang, Bodong
Knudsen, Beatrice S.
Tasdizen, Tolga
Computer Vision and Pattern Recognition
Artificial Intelligence
We here propose a novel hierarchical transformer model that adeptly integrates the feature extraction capabilities of Convolutional Neural Networks (CNNs) with the advanced representational potential of Vision Transformers (ViTs). Addressing the lack of inductive biases and dependence on extensive training datasets in ViTs, our model employs a CNN backbone to generate hierarchical visual representations. These representations are then adapted for transformer input through an innovative patch tokenization. We also introduce a 'scale attention' mechanism that captures cross-scale dependencies, complementing patch attention to enhance spatial understanding and preserve global perception. Our approach significantly outperforms baseline models on small and medium-sized medical datasets, demonstrating its efficiency and generalizability. The components are designed as plug-and-play for different CNN architectures and can be adapted for multiple applications. The code is available at https://github.com/xiaoyatang/DuoFormer.git.
title DuoFormer: Leveraging Hierarchical Visual Representations by Local and Global Attention
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2407.13920