Caterpillar: A Pure-MLP Architecture with Shifted-Pillars-Concatenation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Sun, Jin, Shi, Xiaoshuang, Wang, Zhiyuan, Xu, Kaidi, Shen, Heng Tao, Zhu, Xiaofeng
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917771985551360
author Sun, Jin
Shi, Xiaoshuang
Wang, Zhiyuan
Xu, Kaidi
Shen, Heng Tao
Zhu, Xiaofeng
author_facet Sun, Jin
Shi, Xiaoshuang
Wang, Zhiyuan
Xu, Kaidi
Shen, Heng Tao
Zhu, Xiaofeng
contents Modeling in Computer Vision has evolved to MLPs. Vision MLPs naturally lack local modeling capability, to which the simplest treatment is combined with convolutional layers. Convolution, famous for its sliding window scheme, also suffers from this scheme of redundancy and lower parallel computation. In this paper, we seek to dispense with the windowing scheme and introduce a more elaborate and parallelizable method to exploit locality. To this end, we propose a new MLP module, namely Shifted-Pillars-Concatenation (SPC), that consists of two steps of processes: (1) Pillars-Shift, which generates four neighboring maps by shifting the input image along four directions, and (2) Pillars-Concatenation, which applies linear transformations and concatenation on the maps to aggregate local features. SPC module offers superior local modeling power and performance gains, making it a promising alternative to the convolutional layer. Then, we build a pure-MLP architecture called Caterpillar by replacing the convolutional layer with the SPC module in a hybrid model of sMLPNet. Extensive experiments show Caterpillar's excellent performance on both small-scale and ImageNet-1k classification benchmarks, with remarkable scalability and transfer capability possessed as well. The code is available at https://github.com/sunjin19126/Caterpillar.
format Preprint
id arxiv_https___arxiv_org_abs_2305_17644
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Caterpillar: A Pure-MLP Architecture with Shifted-Pillars-Concatenation
Sun, Jin
Shi, Xiaoshuang
Wang, Zhiyuan
Xu, Kaidi
Shen, Heng Tao
Zhu, Xiaofeng
Computer Vision and Pattern Recognition
Modeling in Computer Vision has evolved to MLPs. Vision MLPs naturally lack local modeling capability, to which the simplest treatment is combined with convolutional layers. Convolution, famous for its sliding window scheme, also suffers from this scheme of redundancy and lower parallel computation. In this paper, we seek to dispense with the windowing scheme and introduce a more elaborate and parallelizable method to exploit locality. To this end, we propose a new MLP module, namely Shifted-Pillars-Concatenation (SPC), that consists of two steps of processes: (1) Pillars-Shift, which generates four neighboring maps by shifting the input image along four directions, and (2) Pillars-Concatenation, which applies linear transformations and concatenation on the maps to aggregate local features. SPC module offers superior local modeling power and performance gains, making it a promising alternative to the convolutional layer. Then, we build a pure-MLP architecture called Caterpillar by replacing the convolutional layer with the SPC module in a hybrid model of sMLPNet. Extensive experiments show Caterpillar's excellent performance on both small-scale and ImageNet-1k classification benchmarks, with remarkable scalability and transfer capability possessed as well. The code is available at https://github.com/sunjin19126/Caterpillar.
title Caterpillar: A Pure-MLP Architecture with Shifted-Pillars-Concatenation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2305.17644