Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Aggarwal, Shubham, Kumar, Lokendra |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
von: Wang, Junxuan, et al.
Veröffentlicht: (2025)
von: Wang, Junxuan, et al.
Veröffentlicht: (2025)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
von: Bozic, Vukasin, et al.
Veröffentlicht: (2023)
von: Bozic, Vukasin, et al.
Veröffentlicht: (2023)
Emergence of Episodic Memory in Transformers: Characterizing Changes in Temporal Structure of Attention Scores During Training
von: Mistry, Deven Mahesh, et al.
Veröffentlicht: (2025)
von: Mistry, Deven Mahesh, et al.
Veröffentlicht: (2025)
Gated Linear Attention Transformers with Hardware-Efficient Training
von: Yang, Songlin, et al.
Veröffentlicht: (2023)
von: Yang, Songlin, et al.
Veröffentlicht: (2023)
Projected Compression: Trainable Projection for Efficient Transformer Compression
von: Stefaniak, Maciej, et al.
Veröffentlicht: (2025)
von: Stefaniak, Maciej, et al.
Veröffentlicht: (2025)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
von: Leemann, Tobias, et al.
Veröffentlicht: (2024)
von: Leemann, Tobias, et al.
Veröffentlicht: (2024)
Understanding Gated Neurons in Transformers from Their Input-Output Functionality
von: Gerstner, Sebastian, et al.
Veröffentlicht: (2025)
von: Gerstner, Sebastian, et al.
Veröffentlicht: (2025)
LASER: Attention with Exponential Transformation
von: Duvvuri, Sai Surya, et al.
Veröffentlicht: (2024)
von: Duvvuri, Sai Surya, et al.
Veröffentlicht: (2024)
BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
von: Lou, Chao, et al.
Veröffentlicht: (2024)
von: Lou, Chao, et al.
Veröffentlicht: (2024)
Generalized Probabilistic Attention Mechanism in Transformers
von: Heo, DongNyeong, et al.
Veröffentlicht: (2024)
von: Heo, DongNyeong, et al.
Veröffentlicht: (2024)
Position-Agnostic Pre-Projection for Transformer Attention: Nonlinear Feature Construction and Content Skip Before Q/K/V
von: Shinde, Chirag
Veröffentlicht: (2026)
von: Shinde, Chirag
Veröffentlicht: (2026)
Beyond Higher Rank: Token-wise Input-Output Projections for Efficient Low-Rank Adaptation
von: Li, Shiwei, et al.
Veröffentlicht: (2025)
von: Li, Shiwei, et al.
Veröffentlicht: (2025)
Learning to Focus: Focal Attention for Selective and Scalable Transformers
von: Ram, Dhananjay, et al.
Veröffentlicht: (2025)
von: Ram, Dhananjay, et al.
Veröffentlicht: (2025)
Improving Transformers with Dynamically Composable Multi-Head Attention
von: Xiao, Da, et al.
Veröffentlicht: (2024)
von: Xiao, Da, et al.
Veröffentlicht: (2024)
Extracting Rule-based Descriptions of Attention Features in Transformers
von: Friedman, Dan, et al.
Veröffentlicht: (2025)
von: Friedman, Dan, et al.
Veröffentlicht: (2025)
Selective Attention Improves Transformer
von: Leviathan, Yaniv, et al.
Veröffentlicht: (2024)
von: Leviathan, Yaniv, et al.
Veröffentlicht: (2024)
Transformer Based Linear Attention with Optimized GPU Kernel Implementation
von: Gerami, Armin, et al.
Veröffentlicht: (2025)
von: Gerami, Armin, et al.
Veröffentlicht: (2025)
SAP: Syntactic Attention Pruning for Transformer-based Language Models
von: Lee, Tzu-Yun, et al.
Veröffentlicht: (2025)
von: Lee, Tzu-Yun, et al.
Veröffentlicht: (2025)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
von: Nawrot, Piotr, et al.
Veröffentlicht: (2025)
von: Nawrot, Piotr, et al.
Veröffentlicht: (2025)
Faster Transformer Decoding: N-gram Masked Self-Attention
von: Chelba, Ciprian, et al.
Veröffentlicht: (2020)
von: Chelba, Ciprian, et al.
Veröffentlicht: (2020)
Selective Attention: Enhancing Transformer through Principled Context Control
von: Zhang, Xuechen, et al.
Veröffentlicht: (2024)
von: Zhang, Xuechen, et al.
Veröffentlicht: (2024)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
von: Musat, Tiberiu
Veröffentlicht: (2024)
von: Musat, Tiberiu
Veröffentlicht: (2024)
Hyper-Connections for Adaptive Multi-Modal MRI Brain Tumor Segmentation
von: Kumar, Lokendra, et al.
Veröffentlicht: (2026)
von: Kumar, Lokendra, et al.
Veröffentlicht: (2026)
RCStat: A Statistical Framework for using Relative Contextualization in Transformers
von: Mahapatra, Debabrata, et al.
Veröffentlicht: (2025)
von: Mahapatra, Debabrata, et al.
Veröffentlicht: (2025)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
von: Mihaila, George
Veröffentlicht: (2026)
von: Mihaila, George
Veröffentlicht: (2026)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
von: Brandon, William, et al.
Veröffentlicht: (2024)
von: Brandon, William, et al.
Veröffentlicht: (2024)
AMPLIFY:Attention-based Mixup for Performance Improvement and Label Smoothing in Transformer
von: Yang, Leixin, et al.
Veröffentlicht: (2023)
von: Yang, Leixin, et al.
Veröffentlicht: (2023)
PaTH Attention: Position Encoding via Accumulating Householder Transformations
von: Yang, Songlin, et al.
Veröffentlicht: (2025)
von: Yang, Songlin, et al.
Veröffentlicht: (2025)
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
von: Dong, Yihe, et al.
Veröffentlicht: (2025)
von: Dong, Yihe, et al.
Veröffentlicht: (2025)
Attention Is Not All You Need: The Importance of Feedforward Networks in Transformer Models
von: Gerber, Isaac
Veröffentlicht: (2025)
von: Gerber, Isaac
Veröffentlicht: (2025)
Learnable Permutation for Structured Sparsity on Transformer Models
von: Li, Zekai, et al.
Veröffentlicht: (2026)
von: Li, Zekai, et al.
Veröffentlicht: (2026)
Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
von: Song, Woomin, et al.
Veröffentlicht: (2025)
von: Song, Woomin, et al.
Veröffentlicht: (2025)
Trainable Transformer in Transformer
von: Panigrahi, Abhishek, et al.
Veröffentlicht: (2023)
von: Panigrahi, Abhishek, et al.
Veröffentlicht: (2023)
Early Transformers: A study on Efficient Training of Transformer Models through Early-Bird Lottery Tickets
von: Cheekati, Shravan
Veröffentlicht: (2024)
von: Cheekati, Shravan
Veröffentlicht: (2024)
Learnable Multi-Scale Wavelet Transformer: A Novel Alternative to Self-Attention
von: Kiruluta, Andrew, et al.
Veröffentlicht: (2025)
von: Kiruluta, Andrew, et al.
Veröffentlicht: (2025)
GLU Attention Improve Transformer
von: Wang, Zehao
Veröffentlicht: (2025)
von: Wang, Zehao
Veröffentlicht: (2025)
What Matters in Transformers? Not All Attention is Needed
von: He, Shwai, et al.
Veröffentlicht: (2024)
von: He, Shwai, et al.
Veröffentlicht: (2024)
Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention
von: Zeris, Athanasios
Veröffentlicht: (2026)
von: Zeris, Athanasios
Veröffentlicht: (2026)
Hadamard Adapter: An Extreme Parameter-Efficient Adapter Tuning Method for Pre-trained Language Models
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
von: Wang, Junxuan, et al.
Veröffentlicht: (2025) -
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
von: Bozic, Vukasin, et al.
Veröffentlicht: (2023) -
Emergence of Episodic Memory in Transformers: Characterizing Changes in Temporal Structure of Attention Scores During Training
von: Mistry, Deven Mahesh, et al.
Veröffentlicht: (2025) -
Gated Linear Attention Transformers with Hardware-Efficient Training
von: Yang, Songlin, et al.
Veröffentlicht: (2023) -
Projected Compression: Trainable Projection for Efficient Transformer Compression
von: Stefaniak, Maciej, et al.
Veröffentlicht: (2025)