Diversity-Guided MLP Reduction for Efficient Large Vision Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Chengchao, Zhu, Hourun, Fang, Gongfan, Wang, Jianxin, Wang, Xinchao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915505567170560
author Shen, Chengchao
Zhu, Hourun
Fang, Gongfan
Wang, Jianxin
Wang, Xinchao
author_facet Shen, Chengchao
Zhu, Hourun
Fang, Gongfan
Wang, Jianxin
Wang, Xinchao
contents Transformer models achieve excellent scaling property, where the performance is improved with the increment of model capacity. However, large-scale model parameters lead to an unaffordable cost of computing and memory. We analyze popular transformer architectures and find that multilayer perceptron (MLP) modules take up the majority of model parameters. To this end, we focus on the recoverability of the compressed models and propose a Diversity-Guided MLP Reduction (DGMR) method to significantly reduce the parameters of large vision transformers with only negligible performance degradation. Specifically, we conduct a Gram-Schmidt weight pruning strategy to eliminate redundant neurons of MLP hidden layer, while preserving weight diversity for better performance recover during distillation. Compared to the model trained from scratch, our pruned model only requires 0.06\% data of LAION-2B (for the training of large vision transformers) without labels (ImageNet-1K) to recover the original performance. Experimental results on several state-of-the-art large vision transformers demonstrate that our method achieves a more than 57.0\% parameter and FLOPs reduction in a near lossless manner. Notably, for EVA-CLIP-E (4.4B), our method accomplishes a 71.5\% parameter and FLOPs reduction without performance degradation. The source code and trained weights are available at https://github.com/visresearch/DGMR.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08591
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Diversity-Guided MLP Reduction for Efficient Large Vision Transformers
Shen, Chengchao
Zhu, Hourun
Fang, Gongfan
Wang, Jianxin
Wang, Xinchao
Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Transformer models achieve excellent scaling property, where the performance is improved with the increment of model capacity. However, large-scale model parameters lead to an unaffordable cost of computing and memory. We analyze popular transformer architectures and find that multilayer perceptron (MLP) modules take up the majority of model parameters. To this end, we focus on the recoverability of the compressed models and propose a Diversity-Guided MLP Reduction (DGMR) method to significantly reduce the parameters of large vision transformers with only negligible performance degradation. Specifically, we conduct a Gram-Schmidt weight pruning strategy to eliminate redundant neurons of MLP hidden layer, while preserving weight diversity for better performance recover during distillation. Compared to the model trained from scratch, our pruned model only requires 0.06\% data of LAION-2B (for the training of large vision transformers) without labels (ImageNet-1K) to recover the original performance. Experimental results on several state-of-the-art large vision transformers demonstrate that our method achieves a more than 57.0\% parameter and FLOPs reduction in a near lossless manner. Notably, for EVA-CLIP-E (4.4B), our method accomplishes a 71.5\% parameter and FLOPs reduction without performance degradation. The source code and trained weights are available at https://github.com/visresearch/DGMR.
title Diversity-Guided MLP Reduction for Efficient Large Vision Transformers
topic Computer Vision and Pattern Recognition
Machine Learning
Multimedia
url https://arxiv.org/abs/2506.08591