Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Amizadeh, Saeed, Abdali, Sara, Li, Yinheng, Koishida, Kazuhito
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916957224173568
author Amizadeh, Saeed
Abdali, Sara
Li, Yinheng
Koishida, Kazuhito
author_facet Amizadeh, Saeed
Abdali, Sara
Li, Yinheng
Koishida, Kazuhito
contents Transformers and their attention mechanism have been revolutionary in the field of Machine Learning. While originally proposed for the language data, they quickly found their way to the image, video, graph, etc. data modalities with various signal geometries. Despite this versatility, generalizing the attention mechanism to scenarios where data is presented at different scales from potentially different modalities is not straightforward. The attempts to incorporate hierarchy and multi-modality within transformers are largely based on ad hoc heuristics, which are not seamlessly generalizable to similar problems with potentially different structures. To address this problem, in this paper, we take a fundamentally different approach: we first propose a mathematical construct to represent multi-modal, multi-scale data. We then mathematically derive the neural attention mechanics for the proposed construct from the first principle of entropy minimization. We show that the derived formulation is optimal in the sense of being the closest to the standard Softmax attention while incorporating the inductive biases originating from the hierarchical/geometric information of the problem. We further propose an efficient algorithm based on dynamic programming to compute our derived attention mechanism. By incorporating it within transformers, we show that the proposed hierarchical attention mechanism not only can be employed to train transformer models in hierarchical/multi-modal settings from scratch, but it can also be used to inject hierarchical information into classical, pre-trained transformer models post training, resulting in more efficient models in zero-shot manner.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15448
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
Amizadeh, Saeed
Abdali, Sara
Li, Yinheng
Koishida, Kazuhito
Machine Learning
Artificial Intelligence
Neural and Evolutionary Computing
Transformers and their attention mechanism have been revolutionary in the field of Machine Learning. While originally proposed for the language data, they quickly found their way to the image, video, graph, etc. data modalities with various signal geometries. Despite this versatility, generalizing the attention mechanism to scenarios where data is presented at different scales from potentially different modalities is not straightforward. The attempts to incorporate hierarchy and multi-modality within transformers are largely based on ad hoc heuristics, which are not seamlessly generalizable to similar problems with potentially different structures. To address this problem, in this paper, we take a fundamentally different approach: we first propose a mathematical construct to represent multi-modal, multi-scale data. We then mathematically derive the neural attention mechanics for the proposed construct from the first principle of entropy minimization. We show that the derived formulation is optimal in the sense of being the closest to the standard Softmax attention while incorporating the inductive biases originating from the hierarchical/geometric information of the problem. We further propose an efficient algorithm based on dynamic programming to compute our derived attention mechanism. By incorporating it within transformers, we show that the proposed hierarchical attention mechanism not only can be employed to train transformer models in hierarchical/multi-modal settings from scratch, but it can also be used to inject hierarchical information into classical, pre-trained transformer models post training, resulting in more efficient models in zero-shot manner.
title Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
topic Machine Learning
Artificial Intelligence
Neural and Evolutionary Computing
url https://arxiv.org/abs/2509.15448