CTA-Net: A CNN-Transformer Aggregation Network for Improving Multi-Scale Feature Extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meng, Chunlei, Yang, Jiacheng, Lin, Wei, Liu, Bowen, Zhang, Hongda, ouyang, chun, Gan, Zhongxue
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913547409162240
author Meng, Chunlei
Yang, Jiacheng
Lin, Wei
Liu, Bowen
Zhang, Hongda
ouyang, chun
Gan, Zhongxue
author_facet Meng, Chunlei
Yang, Jiacheng
Lin, Wei
Liu, Bowen
Zhang, Hongda
ouyang, chun
Gan, Zhongxue
contents Convolutional neural networks (CNNs) and vision transformers (ViTs) have become essential in computer vision for local and global feature extraction. However, aggregating these architectures in existing methods often results in inefficiencies. To address this, the CNN-Transformer Aggregation Network (CTA-Net) was developed. CTA-Net combines CNNs and ViTs, with transformers capturing long-range dependencies and CNNs extracting localized features. This integration enables efficient processing of detailed local and broader contextual information. CTA-Net introduces the Light Weight Multi-Scale Feature Fusion Multi-Head Self-Attention (LMF-MHSA) module for effective multi-scale feature integration with reduced parameters. Additionally, the Reverse Reconstruction CNN-Variants (RRCV) module enhances the embedding of CNNs within the transformer architecture. Extensive experiments on small-scale datasets with fewer than 100,000 samples show that CTA-Net achieves superior performance (TOP-1 Acc 86.76\%), fewer parameters (20.32M), and greater efficiency (FLOPs 2.83B), making it a highly efficient and lightweight solution for visual tasks on small-scale datasets (fewer than 100,000).
format Preprint
id arxiv_https___arxiv_org_abs_2410_11428
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CTA-Net: A CNN-Transformer Aggregation Network for Improving Multi-Scale Feature Extraction
Meng, Chunlei
Yang, Jiacheng
Lin, Wei
Liu, Bowen
Zhang, Hongda
ouyang, chun
Gan, Zhongxue
Computer Vision and Pattern Recognition
Artificial Intelligence
Convolutional neural networks (CNNs) and vision transformers (ViTs) have become essential in computer vision for local and global feature extraction. However, aggregating these architectures in existing methods often results in inefficiencies. To address this, the CNN-Transformer Aggregation Network (CTA-Net) was developed. CTA-Net combines CNNs and ViTs, with transformers capturing long-range dependencies and CNNs extracting localized features. This integration enables efficient processing of detailed local and broader contextual information. CTA-Net introduces the Light Weight Multi-Scale Feature Fusion Multi-Head Self-Attention (LMF-MHSA) module for effective multi-scale feature integration with reduced parameters. Additionally, the Reverse Reconstruction CNN-Variants (RRCV) module enhances the embedding of CNNs within the transformer architecture. Extensive experiments on small-scale datasets with fewer than 100,000 samples show that CTA-Net achieves superior performance (TOP-1 Acc 86.76\%), fewer parameters (20.32M), and greater efficiency (FLOPs 2.83B), making it a highly efficient and lightweight solution for visual tasks on small-scale datasets (fewer than 100,000).
title CTA-Net: A CNN-Transformer Aggregation Network for Improving Multi-Scale Feature Extraction
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2410.11428