Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Aggarwal, Shubham, Kumar, Lokendra |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
di: Wang, Junxuan, et al.
Pubblicazione: (2025)
di: Wang, Junxuan, et al.
Pubblicazione: (2025)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
di: Bozic, Vukasin, et al.
Pubblicazione: (2023)
di: Bozic, Vukasin, et al.
Pubblicazione: (2023)
Emergence of Episodic Memory in Transformers: Characterizing Changes in Temporal Structure of Attention Scores During Training
di: Mistry, Deven Mahesh, et al.
Pubblicazione: (2025)
di: Mistry, Deven Mahesh, et al.
Pubblicazione: (2025)
Gated Linear Attention Transformers with Hardware-Efficient Training
di: Yang, Songlin, et al.
Pubblicazione: (2023)
di: Yang, Songlin, et al.
Pubblicazione: (2023)
Projected Compression: Trainable Projection for Efficient Transformer Compression
di: Stefaniak, Maciej, et al.
Pubblicazione: (2025)
di: Stefaniak, Maciej, et al.
Pubblicazione: (2025)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
di: Leemann, Tobias, et al.
Pubblicazione: (2024)
di: Leemann, Tobias, et al.
Pubblicazione: (2024)
Understanding Gated Neurons in Transformers from Their Input-Output Functionality
di: Gerstner, Sebastian, et al.
Pubblicazione: (2025)
di: Gerstner, Sebastian, et al.
Pubblicazione: (2025)
LASER: Attention with Exponential Transformation
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2024)
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2024)
BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs
di: Wang, Hongyu, et al.
Pubblicazione: (2025)
di: Wang, Hongyu, et al.
Pubblicazione: (2025)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
di: Lou, Chao, et al.
Pubblicazione: (2024)
di: Lou, Chao, et al.
Pubblicazione: (2024)
Generalized Probabilistic Attention Mechanism in Transformers
di: Heo, DongNyeong, et al.
Pubblicazione: (2024)
di: Heo, DongNyeong, et al.
Pubblicazione: (2024)
Position-Agnostic Pre-Projection for Transformer Attention: Nonlinear Feature Construction and Content Skip Before Q/K/V
di: Shinde, Chirag
Pubblicazione: (2026)
di: Shinde, Chirag
Pubblicazione: (2026)
Beyond Higher Rank: Token-wise Input-Output Projections for Efficient Low-Rank Adaptation
di: Li, Shiwei, et al.
Pubblicazione: (2025)
di: Li, Shiwei, et al.
Pubblicazione: (2025)
Learning to Focus: Focal Attention for Selective and Scalable Transformers
di: Ram, Dhananjay, et al.
Pubblicazione: (2025)
di: Ram, Dhananjay, et al.
Pubblicazione: (2025)
Improving Transformers with Dynamically Composable Multi-Head Attention
di: Xiao, Da, et al.
Pubblicazione: (2024)
di: Xiao, Da, et al.
Pubblicazione: (2024)
Extracting Rule-based Descriptions of Attention Features in Transformers
di: Friedman, Dan, et al.
Pubblicazione: (2025)
di: Friedman, Dan, et al.
Pubblicazione: (2025)
Selective Attention Improves Transformer
di: Leviathan, Yaniv, et al.
Pubblicazione: (2024)
di: Leviathan, Yaniv, et al.
Pubblicazione: (2024)
Transformer Based Linear Attention with Optimized GPU Kernel Implementation
di: Gerami, Armin, et al.
Pubblicazione: (2025)
di: Gerami, Armin, et al.
Pubblicazione: (2025)
SAP: Syntactic Attention Pruning for Transformer-based Language Models
di: Lee, Tzu-Yun, et al.
Pubblicazione: (2025)
di: Lee, Tzu-Yun, et al.
Pubblicazione: (2025)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
di: Nawrot, Piotr, et al.
Pubblicazione: (2025)
di: Nawrot, Piotr, et al.
Pubblicazione: (2025)
Faster Transformer Decoding: N-gram Masked Self-Attention
di: Chelba, Ciprian, et al.
Pubblicazione: (2020)
di: Chelba, Ciprian, et al.
Pubblicazione: (2020)
Selective Attention: Enhancing Transformer through Principled Context Control
di: Zhang, Xuechen, et al.
Pubblicazione: (2024)
di: Zhang, Xuechen, et al.
Pubblicazione: (2024)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
di: Musat, Tiberiu
Pubblicazione: (2024)
di: Musat, Tiberiu
Pubblicazione: (2024)
Hyper-Connections for Adaptive Multi-Modal MRI Brain Tumor Segmentation
di: Kumar, Lokendra, et al.
Pubblicazione: (2026)
di: Kumar, Lokendra, et al.
Pubblicazione: (2026)
RCStat: A Statistical Framework for using Relative Contextualization in Transformers
di: Mahapatra, Debabrata, et al.
Pubblicazione: (2025)
di: Mahapatra, Debabrata, et al.
Pubblicazione: (2025)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
di: Mihaila, George
Pubblicazione: (2026)
di: Mihaila, George
Pubblicazione: (2026)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
di: Brandon, William, et al.
Pubblicazione: (2024)
di: Brandon, William, et al.
Pubblicazione: (2024)
AMPLIFY:Attention-based Mixup for Performance Improvement and Label Smoothing in Transformer
di: Yang, Leixin, et al.
Pubblicazione: (2023)
di: Yang, Leixin, et al.
Pubblicazione: (2023)
PaTH Attention: Position Encoding via Accumulating Householder Transformations
di: Yang, Songlin, et al.
Pubblicazione: (2025)
di: Yang, Songlin, et al.
Pubblicazione: (2025)
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
di: Dong, Yihe, et al.
Pubblicazione: (2025)
di: Dong, Yihe, et al.
Pubblicazione: (2025)
Attention Is Not All You Need: The Importance of Feedforward Networks in Transformer Models
di: Gerber, Isaac
Pubblicazione: (2025)
di: Gerber, Isaac
Pubblicazione: (2025)
Learnable Permutation for Structured Sparsity on Transformer Models
di: Li, Zekai, et al.
Pubblicazione: (2026)
di: Li, Zekai, et al.
Pubblicazione: (2026)
Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
di: Song, Woomin, et al.
Pubblicazione: (2025)
di: Song, Woomin, et al.
Pubblicazione: (2025)
Trainable Transformer in Transformer
di: Panigrahi, Abhishek, et al.
Pubblicazione: (2023)
di: Panigrahi, Abhishek, et al.
Pubblicazione: (2023)
Early Transformers: A study on Efficient Training of Transformer Models through Early-Bird Lottery Tickets
di: Cheekati, Shravan
Pubblicazione: (2024)
di: Cheekati, Shravan
Pubblicazione: (2024)
Learnable Multi-Scale Wavelet Transformer: A Novel Alternative to Self-Attention
di: Kiruluta, Andrew, et al.
Pubblicazione: (2025)
di: Kiruluta, Andrew, et al.
Pubblicazione: (2025)
GLU Attention Improve Transformer
di: Wang, Zehao
Pubblicazione: (2025)
di: Wang, Zehao
Pubblicazione: (2025)
What Matters in Transformers? Not All Attention is Needed
di: He, Shwai, et al.
Pubblicazione: (2024)
di: He, Shwai, et al.
Pubblicazione: (2024)
Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention
di: Zeris, Athanasios
Pubblicazione: (2026)
di: Zeris, Athanasios
Pubblicazione: (2026)
Hadamard Adapter: An Extreme Parameter-Efficient Adapter Tuning Method for Pre-trained Language Models
di: Chen, Yuyan, et al.
Pubblicazione: (2024)
di: Chen, Yuyan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
di: Wang, Junxuan, et al.
Pubblicazione: (2025) -
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
di: Bozic, Vukasin, et al.
Pubblicazione: (2023) -
Emergence of Episodic Memory in Transformers: Characterizing Changes in Temporal Structure of Attention Scores During Training
di: Mistry, Deven Mahesh, et al.
Pubblicazione: (2025) -
Gated Linear Attention Transformers with Hardware-Efficient Training
di: Yang, Songlin, et al.
Pubblicazione: (2023) -
Projected Compression: Trainable Projection for Efficient Transformer Compression
di: Stefaniak, Maciej, et al.
Pubblicazione: (2025)