Saccade Attention Networks: Using Transfer Learning of Attention to Reduce Network Sizes

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteur principal: Estafanous, Marc
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914485228273664
author Estafanous, Marc
author_facet Estafanous, Marc
contents One of the limitations of transformer networks is the sequence length due to the quadratic nature of the attention matrix. Classical self attention uses the entire sequence length, however, the actual attention being used is sparse. Humans use a form of sparse attention when analyzing an image or scene called saccades. Focusing on key features greatly reduces computation time. By using a network (Saccade Attention Network) to learn where to attend from a large pre-trained model, we can use it to pre-process images and greatly reduce network size by reducing the input sequence length to just the key features being attended to. Our results indicate that you can reduce calculations by close to 80% and produce similar results.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16485
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Saccade Attention Networks: Using Transfer Learning of Attention to Reduce Network Sizes
Estafanous, Marc
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
I.4; I.5.1
One of the limitations of transformer networks is the sequence length due to the quadratic nature of the attention matrix. Classical self attention uses the entire sequence length, however, the actual attention being used is sparse. Humans use a form of sparse attention when analyzing an image or scene called saccades. Focusing on key features greatly reduces computation time. By using a network (Saccade Attention Network) to learn where to attend from a large pre-trained model, we can use it to pre-process images and greatly reduce network size by reducing the input sequence length to just the key features being attended to. Our results indicate that you can reduce calculations by close to 80% and produce similar results.
title Saccade Attention Networks: Using Transfer Learning of Attention to Reduce Network Sizes
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
I.4; I.5.1
url https://arxiv.org/abs/2604.16485