Enhancing Layer Attention Efficiency through Pruning Redundant Retrievals

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Hanze, Du, Yaosong, Yao, Zhibo, Zeng, Mengyao, Ge, Xiuqi, Huang, Xiande
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916069337202688
author Li, Hanze
Du, Yaosong
Yao, Zhibo
Zeng, Mengyao
Ge, Xiuqi
Huang, Xiande
author_facet Li, Hanze
Du, Yaosong
Yao, Zhibo
Zeng, Mengyao
Ge, Xiuqi
Huang, Xiande
contents Growing evidence suggests that layer attention mechanisms, which enhance interaction among layers in deep neural networks, have significantly advanced network architectures. However, existing layer attention methods suffer from redundancy, as attention weights learned by adjacent layers often become highly similar. This redundancy causes multiple layers to extract nearly identical features, reducing the model's representational capacity and increasing training time. To address this issue, we propose a novel approach to quantify redundancy by leveraging the Kullback-Leibler (KL) divergence between adjacent layers. Additionally, we introduce an Enhanced Beta Quantile Mapping (EBQM) method that accurately identifies and skips redundant layers, thereby maintaining model stability. Our proposed Efficient Layer Attention (ELA) architecture, improves both training efficiency and overall performance, achieving a 30% reduction in training time while enhancing performance in tasks such as image classification and object detection.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06473
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Layer Attention Efficiency through Pruning Redundant Retrievals
Li, Hanze
Du, Yaosong
Yao, Zhibo
Zeng, Mengyao
Ge, Xiuqi
Huang, Xiande
Computer Vision and Pattern Recognition
Artificial Intelligence
Growing evidence suggests that layer attention mechanisms, which enhance interaction among layers in deep neural networks, have significantly advanced network architectures. However, existing layer attention methods suffer from redundancy, as attention weights learned by adjacent layers often become highly similar. This redundancy causes multiple layers to extract nearly identical features, reducing the model's representational capacity and increasing training time. To address this issue, we propose a novel approach to quantify redundancy by leveraging the Kullback-Leibler (KL) divergence between adjacent layers. Additionally, we introduce an Enhanced Beta Quantile Mapping (EBQM) method that accurately identifies and skips redundant layers, thereby maintaining model stability. Our proposed Efficient Layer Attention (ELA) architecture, improves both training efficiency and overall performance, achieving a 30% reduction in training time while enhancing performance in tasks such as image classification and object detection.
title Enhancing Layer Attention Efficiency through Pruning Redundant Retrievals
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.06473