Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
Fuente:
arXiv
Saved in:
| Main Authors: | Amizadeh, Saeed, Abdali, Sara, Li, Yinheng, Koishida, Kazuhito |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attentions Under the Microscope: A Comparative Study of Resource Utilization for Variants of Self-Attention
by: Tian, Zhengyu, et al.
Published: (2025)
by: Tian, Zhengyu, et al.
Published: (2025)
Enhancing Graph Representation Learning with Attention-Driven Spiking Neural Networks
by: Yin, Huifeng, et al.
Published: (2024)
by: Yin, Huifeng, et al.
Published: (2024)
SemlaFlow -- Efficient 3D Molecular Generation with Latent Attention and Equivariant Flow Matching
by: Irwin, Ross, et al.
Published: (2024)
by: Irwin, Ross, et al.
Published: (2024)
Greener GRASS: Enhancing GNNs with Encoding, Rewiring, and Attention
by: Liao, Tongzhou, et al.
Published: (2024)
by: Liao, Tongzhou, et al.
Published: (2024)
Neural Dynamics Self-Attention for Spiking Transformers
by: Zhang, Dehao, et al.
Published: (2026)
by: Zhang, Dehao, et al.
Published: (2026)
Beyond Attention: Toward Machines with Intrinsic Higher Mental States
by: Adeel, Ahsan
Published: (2025)
by: Adeel, Ahsan
Published: (2025)
Breaking Global Self-Attention Bottlenecks in Transformer-based Spiking Neural Networks with Local Structure-Aware Self-Attention
by: Li, Lingdong, et al.
Published: (2026)
by: Li, Lingdong, et al.
Published: (2026)
Hierarchical Kernel Transformer: Multi-Scale Attention with an Information-Theoretic Approximation Analysis
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
Sparsity Moves Computation: How FFN Architecture Reshapes Attention in Small Transformers
by: Smithline, Gabriel, et al.
Published: (2026)
by: Smithline, Gabriel, et al.
Published: (2026)
Backpropagation-free Spiking Neural Networks with the Forward-Forward Algorithm
by: Ghader, Mohammadnavid, et al.
Published: (2025)
by: Ghader, Mohammadnavid, et al.
Published: (2025)
Selective Synchronization Attention
by: Hays, Hasi
Published: (2026)
by: Hays, Hasi
Published: (2026)
GLU Attention Improve Transformer
by: Wang, Zehao
Published: (2025)
by: Wang, Zehao
Published: (2025)
A Resource Model For Neural Scaling Law
by: Song, Jinyeop, et al.
Published: (2024)
by: Song, Jinyeop, et al.
Published: (2024)
Deep ARTMAP: Generalized Hierarchical Learning with Adaptive Resonance Theory
by: Melton, Niklas M., et al.
Published: (2025)
by: Melton, Niklas M., et al.
Published: (2025)
What Planning Problems Can A Relational Neural Network Solve?
by: Mao, Jiayuan, et al.
Published: (2023)
by: Mao, Jiayuan, et al.
Published: (2023)
Beyond Uniform Scaling: Exploring Depth Heterogeneity in Neural Architectures
by: T, Akash Guna R., et al.
Published: (2024)
by: T, Akash Guna R., et al.
Published: (2024)
BiSpikCLM: A Spiking Language Model integrating Softmax-Free Spiking Attention and Spike-Aware Alignment Distillation
by: Guo, Sihang, et al.
Published: (2026)
by: Guo, Sihang, et al.
Published: (2026)
AutoQD: Automatic Discovery of Diverse Behaviors with Quality-Diversity Optimization
by: Hedayatian, Saeed, et al.
Published: (2025)
by: Hedayatian, Saeed, et al.
Published: (2025)
Feedback Favors the Generalization of Neural ODEs
by: Jia, Jindou, et al.
Published: (2024)
by: Jia, Jindou, et al.
Published: (2024)
Bridging Synthetic and Real Routing Problems via LLM-Guided Instance Generation and Progressive Adaptation
by: Zhu, Jianghan, et al.
Published: (2025)
by: Zhu, Jianghan, et al.
Published: (2025)
Towards evolution of Deep Neural Networks through contrastive Self-Supervised learning
by: Vinhas, Adriano, et al.
Published: (2024)
by: Vinhas, Adriano, et al.
Published: (2024)
Multi-scale Topology Optimization using Neural Networks
by: Chen, Hongrui, et al.
Published: (2024)
by: Chen, Hongrui, et al.
Published: (2024)
Preisach Attention: A Hysteretic Model of Sequential Memory
by: Frydrych, Piotr
Published: (2026)
by: Frydrych, Piotr
Published: (2026)
Quantum-Evolutionary Neural Networks for Multi-Agent Federated Learning
by: Lala, Aarav, et al.
Published: (2025)
by: Lala, Aarav, et al.
Published: (2025)
Evolution Meets Diffusion: Efficient Neural Architecture Generation
by: Zhou, Bingye, et al.
Published: (2025)
by: Zhou, Bingye, et al.
Published: (2025)
Structure as Computation: Developmental Generation of Minimal Neural Circuits
by: Zhou, Duan
Published: (2026)
by: Zhou, Duan
Published: (2026)
DGPO: RL-Steered Graph Diffusion for Neural Architecture Generation
by: Liuliakov, Aleksei, et al.
Published: (2026)
by: Liuliakov, Aleksei, et al.
Published: (2026)
MTSpark: Enabling Multi-Task Learning with Spiking Neural Networks for Generalist Agents
by: Devkota, Avaneesh, et al.
Published: (2024)
by: Devkota, Avaneesh, et al.
Published: (2024)
SAN: Hypothesizing Long-Term Synaptic Development and Neural Engram Mechanism in Scalable Model's Parameter-Efficient Fine-Tuning
by: Dai, Gaole, et al.
Published: (2024)
by: Dai, Gaole, et al.
Published: (2024)
Structurally Flexible Neural Networks: Evolving the Building Blocks for General Agents
by: Pedersen, Joachim Winther, et al.
Published: (2024)
by: Pedersen, Joachim Winther, et al.
Published: (2024)
On the Improvement of Generalization and Stability of Forward-Only Learning via Neural Polarization
by: Terres-Escudero, Erik B., et al.
Published: (2024)
by: Terres-Escudero, Erik B., et al.
Published: (2024)
Instance-Conditioned Adaptation for Large-scale Generalization of Neural Routing Solver
by: Zhou, Changliang, et al.
Published: (2024)
by: Zhou, Changliang, et al.
Published: (2024)
Hierarchical Residuals Exploit Brain-Inspired Compositionality
by: López, Francisco M., et al.
Published: (2025)
by: López, Francisco M., et al.
Published: (2025)
Application of Unsupervised Artificial Neural Network (ANN) Self_Organizing Map (SOM) in Identifying Main Car Sales Factors
by: Taghavi, Mazyar
Published: (2024)
by: Taghavi, Mazyar
Published: (2024)
The Ungrounded Alignment Problem
by: Pickett, Marc, et al.
Published: (2024)
by: Pickett, Marc, et al.
Published: (2024)
Continuous Spiking Graph Neural Networks
by: Yin, Nan, et al.
Published: (2024)
by: Yin, Nan, et al.
Published: (2024)
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
by: Csordás, Róbert, et al.
Published: (2024)
by: Csordás, Róbert, et al.
Published: (2024)
ActPC-Geom: Towards Scalable Online Neural-Symbolic Learning via Accelerating Active Predictive Coding with Information Geometry & Diverse Cognitive Mechanisms
by: Goertzel, Ben
Published: (2025)
by: Goertzel, Ben
Published: (2025)
SGNNBench: A Holistic Evaluation of Spiking Graph Neural Network on Large-scale Graph
by: Zhang, Huizhe, et al.
Published: (2025)
by: Zhang, Huizhe, et al.
Published: (2025)
Sample-based Dynamic Hierarchical Transformer with Layer and Head Flexibility via Contextual Bandit
by: Meng, Fanfei, et al.
Published: (2023)
by: Meng, Fanfei, et al.
Published: (2023)
Similar Items
-
Attentions Under the Microscope: A Comparative Study of Resource Utilization for Variants of Self-Attention
by: Tian, Zhengyu, et al.
Published: (2025) -
Enhancing Graph Representation Learning with Attention-Driven Spiking Neural Networks
by: Yin, Huifeng, et al.
Published: (2024) -
SemlaFlow -- Efficient 3D Molecular Generation with Latent Attention and Equivariant Flow Matching
by: Irwin, Ross, et al.
Published: (2024) -
Greener GRASS: Enhancing GNNs with Encoding, Rewiring, and Attention
by: Liao, Tongzhou, et al.
Published: (2024) -
Neural Dynamics Self-Attention for Spiking Transformers
by: Zhang, Dehao, et al.
Published: (2026)