Convolutional Rectangular Attention Module

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Hai-Vy, Gamboa, Fabrice, Zhang, Sixin, Chhaibi, Reda, Gratton, Serge, Giaccone, Thierry
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909759026757632
author Nguyen, Hai-Vy
Gamboa, Fabrice
Zhang, Sixin
Chhaibi, Reda
Gratton, Serge
Giaccone, Thierry
author_facet Nguyen, Hai-Vy
Gamboa, Fabrice
Zhang, Sixin
Chhaibi, Reda
Gratton, Serge
Giaccone, Thierry
contents In this paper, we introduce a novel spatial attention module that can be easily integrated to any convolutional network. This module guides the model to pay attention to the most discriminative part of an image. This enables the model to attain a better performance by an end-to-end training. In conventional approaches, a spatial attention map is typically generated in a position-wise manner. Thus, it is often resulting in irregular boundaries and so can hamper generalization to new samples. In our method, the attention region is constrained to be rectangular. This rectangle is parametrized by only 5 parameters, allowing for a better stability and generalization to new samples. In our experiments, our method systematically outperforms the position-wise counterpart. So that, we provide a novel useful spatial attention mechanism for convolutional models. Besides, our module also provides the interpretability regarding the \textit{where to look} question, as it helps to know the part of the input on which the model focuses to produce the prediction.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10875
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Convolutional Rectangular Attention Module
Nguyen, Hai-Vy
Gamboa, Fabrice
Zhang, Sixin
Chhaibi, Reda
Gratton, Serge
Giaccone, Thierry
Computer Vision and Pattern Recognition
Machine Learning
In this paper, we introduce a novel spatial attention module that can be easily integrated to any convolutional network. This module guides the model to pay attention to the most discriminative part of an image. This enables the model to attain a better performance by an end-to-end training. In conventional approaches, a spatial attention map is typically generated in a position-wise manner. Thus, it is often resulting in irregular boundaries and so can hamper generalization to new samples. In our method, the attention region is constrained to be rectangular. This rectangle is parametrized by only 5 parameters, allowing for a better stability and generalization to new samples. In our experiments, our method systematically outperforms the position-wise counterpart. So that, we provide a novel useful spatial attention mechanism for convolutional models. Besides, our module also provides the interpretability regarding the \textit{where to look} question, as it helps to know the part of the input on which the model focuses to produce the prediction.
title Convolutional Rectangular Attention Module
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2503.10875