On the Role of Attention Masks and LayerNorm in Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Xinyi, Ajorlou, Amir, Wang, Yifei, Jegelka, Stefanie, Jadbabaie, Ali |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Emergence of Position Bias in Transformers
by: Wu, Xinyi, et al.
Published: (2025)
by: Wu, Xinyi, et al.
Published: (2025)
Demystifying Oversmoothing in Attention-Based Graph Neural Networks
by: Wu, Xinyi, et al.
Published: (2023)
by: Wu, Xinyi, et al.
Published: (2023)
Geometry and Dynamics of LayerNorm
by: Riechers, Paul M.
Published: (2024)
by: Riechers, Paul M.
Published: (2024)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
by: Baroni, Luca, et al.
Published: (2025)
by: Baroni, Luca, et al.
Published: (2025)
LayerNorm Induces Recency Bias in Transformer Decoders
by: Kim, Junu, et al.
Published: (2025)
by: Kim, Junu, et al.
Published: (2025)
Belief Samples Are All You Need For Social Learning
by: JafariNodeh, Mahyar, et al.
Published: (2024)
by: JafariNodeh, Mahyar, et al.
Published: (2024)
SLaNC: Static LayerNorm Calibration
by: Salmani, Mahsa, et al.
Published: (2024)
by: Salmani, Mahsa, et al.
Published: (2024)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
by: Wang, Wenxun, et al.
Published: (2025)
by: Wang, Wenxun, et al.
Published: (2025)
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
On the Role of Transformer Feed-Forward Layers in Nonlinear In-Context Learning
by: Sun, Haoyuan, et al.
Published: (2025)
by: Sun, Haoyuan, et al.
Published: (2025)
You can remove GPT2's LayerNorm by fine-tuning
by: Heimersheim, Stefan
Published: (2024)
by: Heimersheim, Stefan
Published: (2024)
LayerNorm: A key component in parameter-efficient fine-tuning
by: ValizadehAslani, Taha, et al.
Published: (2024)
by: ValizadehAslani, Taha, et al.
Published: (2024)
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
by: Qian, Wenhao, et al.
Published: (2025)
by: Qian, Wenhao, et al.
Published: (2025)
When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer
by: Verma, Lucky
Published: (2026)
by: Verma, Lucky
Published: (2026)
Residual Connections and Normalization Can Provably Prevent Oversmoothing in GNNs
by: Scholkemper, Michael, et al.
Published: (2024)
by: Scholkemper, Michael, et al.
Published: (2024)
MoETTA: Test-Time Adaptation Under Mixed Distribution Shifts with MoE-LayerNorm
by: Fan, Xiao, et al.
Published: (2025)
by: Fan, Xiao, et al.
Published: (2025)
A Theoretical Understanding of Self-Correction through In-context Alignment
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Scaling Attention via Feature Sparsity
by: Xie, Yan, et al.
Published: (2026)
by: Xie, Yan, et al.
Published: (2026)
How to Craft Backdoors with Unlabeled Data Alone?
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
A Canonicalization Perspective on Invariant and Equivariant Learning
by: Ma, George, et al.
Published: (2024)
by: Ma, George, et al.
Published: (2024)
Higher-Order Graphon Neural Networks: Approximation and Cut Distance
by: Herbst, Daniel, et al.
Published: (2025)
by: Herbst, Daniel, et al.
Published: (2025)
Sample Complexity Bounds for Estimating Probability Divergences under Invariances
by: Tahmasebi, Behrooz, et al.
Published: (2023)
by: Tahmasebi, Behrooz, et al.
Published: (2023)
The Exact Sample Complexity Gain from Invariances for Kernel Regression
by: Tahmasebi, Behrooz, et al.
Published: (2023)
by: Tahmasebi, Behrooz, et al.
Published: (2023)
LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning
by: Guo, Xiaojun, et al.
Published: (2025)
by: Guo, Xiaojun, et al.
Published: (2025)
Understanding the Role of Equivariance in Self-supervised Learning
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Estimating True Beliefs in Opinion Dynamics with Social Pressure
by: Tang, Jennifer, et al.
Published: (2023)
by: Tang, Jennifer, et al.
Published: (2023)
Stochastic Opinion Dynamics under Social Pressure in Arbitrary Networks
by: Tang, Jennifer, et al.
Published: (2023)
by: Tang, Jennifer, et al.
Published: (2023)
In-Context Symmetries: Self-Supervised Learning through Contextual World Models
by: Gupta, Sharut, et al.
Published: (2024)
by: Gupta, Sharut, et al.
Published: (2024)
When More is Less: Understanding Chain-of-Thought Length in LLMs
by: Wu, Yuyang, et al.
Published: (2025)
by: Wu, Yuyang, et al.
Published: (2025)
Generalization, Expressivity, and Universality of Graph Neural Networks on Attributed Graphs
by: Rauchwerger, Levi, et al.
Published: (2024)
by: Rauchwerger, Levi, et al.
Published: (2024)
Neural Networks With Dense Weights Are Not Universal Approximators
by: Rauchwerger, Levi, et al.
Published: (2026)
by: Rauchwerger, Levi, et al.
Published: (2026)
Counting Substructures with Higher-Order Graph Neural Networks: Possibility and Impossibility Results
by: Tahmasebi, Behrooz, et al.
Published: (2020)
by: Tahmasebi, Behrooz, et al.
Published: (2020)
A Poincaré Inequality and Consistency Results for Signal Sampling on Large Graphs
by: Le, Thien, et al.
Published: (2023)
by: Le, Thien, et al.
Published: (2023)
Masked BRep Autoencoder via Hierarchical Graph Transformer
by: Li, Yifei, et al.
Published: (2026)
by: Li, Yifei, et al.
Published: (2026)
A least-square method for non-asymptotic identification in linear switching control
by: Sun, Haoyuan, et al.
Published: (2024)
by: Sun, Haoyuan, et al.
Published: (2024)
Online Learning for Supervisory Switching Control
by: Sun, Haoyuan, et al.
Published: (2026)
by: Sun, Haoyuan, et al.
Published: (2026)
Convolutional Learning on Directed Acyclic Graphs
by: Rey, Samuel, et al.
Published: (2024)
by: Rey, Samuel, et al.
Published: (2024)
On the Role of Depth and Looping for In-Context Learning with Task Diversity
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval
by: Guo, Lixuan, et al.
Published: (2026)
by: Guo, Lixuan, et al.
Published: (2026)
Similar Items
-
On the Emergence of Position Bias in Transformers
by: Wu, Xinyi, et al.
Published: (2025) -
Demystifying Oversmoothing in Attention-Based Graph Neural Networks
by: Wu, Xinyi, et al.
Published: (2023) -
Geometry and Dynamics of LayerNorm
by: Riechers, Paul M.
Published: (2024) -
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
by: Baroni, Luca, et al.
Published: (2025) -
LayerNorm Induces Recency Bias in Transformer Decoders
by: Kim, Junu, et al.
Published: (2025)