Static Key Attention in Vision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Zizhao, Zhou, Xiaolin, Rostami, Mohammad
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929621422833664
author Hu, Zizhao
Zhou, Xiaolin
Rostami, Mohammad
author_facet Hu, Zizhao
Zhou, Xiaolin
Rostami, Mohammad
contents The success of vision transformers is widely attributed to the expressive power of their dynamically parameterized multi-head self-attention mechanism. We examine the impact of substituting the dynamic parameterized key with a static key within the standard attention mechanism in Vision Transformers. Our findings reveal that static key attention mechanisms can match or even exceed the performance of standard self-attention. Integrating static key attention modules into a Metaformer backbone, we find that it serves as a better intermediate stage in hierarchical hybrid architectures, balancing the strengths of depth-wise convolution and self-attention. Experiments on several vision tasks underscore the effectiveness of the static key mechanism, indicating that the typical two-step dynamic parameterization in attention can be streamlined to a single step without impacting performance under certain circumstances.
format Preprint
id arxiv_https___arxiv_org_abs_2412_07049
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Static Key Attention in Vision
Hu, Zizhao
Zhou, Xiaolin
Rostami, Mohammad
Computer Vision and Pattern Recognition
The success of vision transformers is widely attributed to the expressive power of their dynamically parameterized multi-head self-attention mechanism. We examine the impact of substituting the dynamic parameterized key with a static key within the standard attention mechanism in Vision Transformers. Our findings reveal that static key attention mechanisms can match or even exceed the performance of standard self-attention. Integrating static key attention modules into a Metaformer backbone, we find that it serves as a better intermediate stage in hierarchical hybrid architectures, balancing the strengths of depth-wise convolution and self-attention. Experiments on several vision tasks underscore the effectiveness of the static key mechanism, indicating that the typical two-step dynamic parameterization in attention can be streamlined to a single step without impacting performance under certain circumstances.
title Static Key Attention in Vision
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.07049