Sparsity Moves Computation: How FFN Architecture Reshapes Attention in Small Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Smithline, Gabriel, Mascioli, Chris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Weight Sparsity Complements Activity Sparsity in Neuromorphic Language Models
von: Mukherji, Rishav, et al.
Veröffentlicht: (2024)
von: Mukherji, Rishav, et al.
Veröffentlicht: (2024)
CLASSP: a Biologically-Inspired Approach to Continual Learning through Adjustment Suppression and Sparsity Promotion
von: Ludwig, Oswaldo
Veröffentlicht: (2024)
von: Ludwig, Oswaldo
Veröffentlicht: (2024)
Explore Activation Sparsity in Recurrent LLMs for Energy-Efficient Neuromorphic Computing
von: Knunyants, Ivan, et al.
Veröffentlicht: (2025)
von: Knunyants, Ivan, et al.
Veröffentlicht: (2025)
GLU Attention Improve Transformer
von: Wang, Zehao
Veröffentlicht: (2025)
von: Wang, Zehao
Veröffentlicht: (2025)
Sup3r: A Semi-Supervised Algorithm for increasing Sparsity, Stability, and Separability in Hierarchy Of Time-Surfaces architectures
von: Rasetto, Marco, et al.
Veröffentlicht: (2024)
von: Rasetto, Marco, et al.
Veröffentlicht: (2024)
Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
von: Amizadeh, Saeed, et al.
Veröffentlicht: (2025)
von: Amizadeh, Saeed, et al.
Veröffentlicht: (2025)
Small Contributions, Small Networks: Efficient Neural Network Pruning Based on Relative Importance
von: Hussien, Mostafa, et al.
Veröffentlicht: (2024)
von: Hussien, Mostafa, et al.
Veröffentlicht: (2024)
Attentions Under the Microscope: A Comparative Study of Resource Utilization for Variants of Self-Attention
von: Tian, Zhengyu, et al.
Veröffentlicht: (2025)
von: Tian, Zhengyu, et al.
Veröffentlicht: (2025)
STAR: Synthesis of Tailored Architectures
von: Thomas, Armin W., et al.
Veröffentlicht: (2024)
von: Thomas, Armin W., et al.
Veröffentlicht: (2024)
A Transformer-based Neural Architecture Search Method
von: Wang, Shang, et al.
Veröffentlicht: (2025)
von: Wang, Shang, et al.
Veröffentlicht: (2025)
MIDAS: Mosaic Input-Specific Differentiable Architecture Search
von: Subbotko, Konstanty
Veröffentlicht: (2026)
von: Subbotko, Konstanty
Veröffentlicht: (2026)
Evolution Meets Diffusion: Efficient Neural Architecture Generation
von: Zhou, Bingye, et al.
Veröffentlicht: (2025)
von: Zhou, Bingye, et al.
Veröffentlicht: (2025)
Knowledge-aware Evolutionary Graph Neural Architecture Search
von: Wang, Chao, et al.
Veröffentlicht: (2024)
von: Wang, Chao, et al.
Veröffentlicht: (2024)
Greener GRASS: Enhancing GNNs with Encoding, Rewiring, and Attention
von: Liao, Tongzhou, et al.
Veröffentlicht: (2024)
von: Liao, Tongzhou, et al.
Veröffentlicht: (2024)
DGPO: RL-Steered Graph Diffusion for Neural Architecture Generation
von: Liuliakov, Aleksei, et al.
Veröffentlicht: (2026)
von: Liuliakov, Aleksei, et al.
Veröffentlicht: (2026)
Beyond Uniform Scaling: Exploring Depth Heterogeneity in Neural Architectures
von: T, Akash Guna R., et al.
Veröffentlicht: (2024)
von: T, Akash Guna R., et al.
Veröffentlicht: (2024)
Evolutionary Architecture Search through Grammar-Based Sequence Alignment
von: Martín, Adri Gómez, et al.
Veröffentlicht: (2025)
von: Martín, Adri Gómez, et al.
Veröffentlicht: (2025)
Beyond Attention: Toward Machines with Intrinsic Higher Mental States
von: Adeel, Ahsan
Veröffentlicht: (2025)
von: Adeel, Ahsan
Veröffentlicht: (2025)
Attending to Graph Transformers
von: Müller, Luis, et al.
Veröffentlicht: (2023)
von: Müller, Luis, et al.
Veröffentlicht: (2023)
SEval-NAS: A Search-Agnostic Evaluation for Neural Architecture Search
von: Mih, Atah Nuh, et al.
Veröffentlicht: (2026)
von: Mih, Atah Nuh, et al.
Veröffentlicht: (2026)
Neural Architecture Search using Particle Swarm and Ant Colony Optimization
von: Lankford, Séamus, et al.
Veröffentlicht: (2024)
von: Lankford, Séamus, et al.
Veröffentlicht: (2024)
Enhancing Graph Representation Learning with Attention-Driven Spiking Neural Networks
von: Yin, Huifeng, et al.
Veröffentlicht: (2024)
von: Yin, Huifeng, et al.
Veröffentlicht: (2024)
Memoria: Resolving Fateful Forgetting Problem through Human-Inspired Memory Architecture
von: Park, Sangjun, et al.
Veröffentlicht: (2023)
von: Park, Sangjun, et al.
Veröffentlicht: (2023)
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2024)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2024)
Structure Development in List-Sorting Transformers
von: Urdshals, Einar, et al.
Veröffentlicht: (2025)
von: Urdshals, Einar, et al.
Veröffentlicht: (2025)
Investigating Recurrent Transformers with Dynamic Halt
von: Chowdhury, Jishnu Ray, et al.
Veröffentlicht: (2024)
von: Chowdhury, Jishnu Ray, et al.
Veröffentlicht: (2024)
Experimental Evaluation of Machine Learning Models for Goal-oriented Customer Service Chatbot with Pipeline Architecture
von: Isa, Nurul Ain Nabilah Mohd, et al.
Veröffentlicht: (2024)
von: Isa, Nurul Ain Nabilah Mohd, et al.
Veröffentlicht: (2024)
Understanding Transformer Optimization via Gradient Heterogeneity
von: Tomihari, Akiyoshi, et al.
Veröffentlicht: (2025)
von: Tomihari, Akiyoshi, et al.
Veröffentlicht: (2025)
Spiking Point Transformer for Point Cloud Classification
von: Wu, Peixi, et al.
Veröffentlicht: (2025)
von: Wu, Peixi, et al.
Veröffentlicht: (2025)
MoEUT: Mixture-of-Experts Universal Transformers
von: Csordás, Róbert, et al.
Veröffentlicht: (2024)
von: Csordás, Róbert, et al.
Veröffentlicht: (2024)
SemlaFlow -- Efficient 3D Molecular Generation with Latent Attention and Equivariant Flow Matching
von: Irwin, Ross, et al.
Veröffentlicht: (2024)
von: Irwin, Ross, et al.
Veröffentlicht: (2024)
High Performance Im2win and Direct Convolutions using Three Tensor Layouts on SIMD Architectures
von: Fu, Xiang, et al.
Veröffentlicht: (2024)
von: Fu, Xiang, et al.
Veröffentlicht: (2024)
General-Purpose In-Context Learning by Meta-Learning Transformers
von: Kirsch, Louis, et al.
Veröffentlicht: (2022)
von: Kirsch, Louis, et al.
Veröffentlicht: (2022)
QSViT: A Methodology for Quantizing Spiking Vision Transformers
von: Putra, Rachmad Vidya Wicaksana, et al.
Veröffentlicht: (2025)
von: Putra, Rachmad Vidya Wicaksana, et al.
Veröffentlicht: (2025)
SGHormer: An Energy-Saving Graph Transformer Driven by Spikes
von: Zhang, Huizhe, et al.
Veröffentlicht: (2024)
von: Zhang, Huizhe, et al.
Veröffentlicht: (2024)
Structure as Computation: Developmental Generation of Minimal Neural Circuits
von: Zhou, Duan
Veröffentlicht: (2026)
von: Zhou, Duan
Veröffentlicht: (2026)
Brain-inspired Computational Intelligence via Predictive Coding
von: Salvatori, Tommaso, et al.
Veröffentlicht: (2023)
von: Salvatori, Tommaso, et al.
Veröffentlicht: (2023)
Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics
von: Verma, Lucky
Veröffentlicht: (2026)
von: Verma, Lucky
Veröffentlicht: (2026)
Decision SpikeFormer: Spike-Driven Transformer for Decision Making
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain
von: Kosowski, Adrian, et al.
Veröffentlicht: (2025)
von: Kosowski, Adrian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Weight Sparsity Complements Activity Sparsity in Neuromorphic Language Models
von: Mukherji, Rishav, et al.
Veröffentlicht: (2024) -
CLASSP: a Biologically-Inspired Approach to Continual Learning through Adjustment Suppression and Sparsity Promotion
von: Ludwig, Oswaldo
Veröffentlicht: (2024) -
Explore Activation Sparsity in Recurrent LLMs for Energy-Efficient Neuromorphic Computing
von: Knunyants, Ivan, et al.
Veröffentlicht: (2025) -
GLU Attention Improve Transformer
von: Wang, Zehao
Veröffentlicht: (2025) -
Sup3r: A Semi-Supervised Algorithm for increasing Sparsity, Stability, and Separability in Hierarchy Of Time-Surfaces architectures
von: Rasetto, Marco, et al.
Veröffentlicht: (2024)