When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer
Fuente:
arXiv
Saved in:
| Main Author: | Verma, Lucky |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LayerNorm Induces Recency Bias in Transformer Decoders
by: Kim, Junu, et al.
Published: (2025)
by: Kim, Junu, et al.
Published: (2025)
SLaNC: Static LayerNorm Calibration
by: Salmani, Mahsa, et al.
Published: (2024)
by: Salmani, Mahsa, et al.
Published: (2024)
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
You can remove GPT2's LayerNorm by fine-tuning
by: Heimersheim, Stefan
Published: (2024)
by: Heimersheim, Stefan
Published: (2024)
LayerNorm: A key component in parameter-efficient fine-tuning
by: ValizadehAslani, Taha, et al.
Published: (2024)
by: ValizadehAslani, Taha, et al.
Published: (2024)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
by: Baroni, Luca, et al.
Published: (2025)
by: Baroni, Luca, et al.
Published: (2025)
Geometry and Dynamics of LayerNorm
by: Riechers, Paul M.
Published: (2024)
by: Riechers, Paul M.
Published: (2024)
On the Role of Attention Masks and LayerNorm in Transformers
by: Wu, Xinyi, et al.
Published: (2024)
by: Wu, Xinyi, et al.
Published: (2024)
Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics
by: Verma, Lucky
Published: (2026)
by: Verma, Lucky
Published: (2026)
When Do "More Contexts" Help with Sarcasm Recognition?
by: Nimase, Ojas, et al.
Published: (2024)
by: Nimase, Ojas, et al.
Published: (2024)
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
by: Qian, Wenhao, et al.
Published: (2025)
by: Qian, Wenhao, et al.
Published: (2025)
MoETTA: Test-Time Adaptation Under Mixed Distribution Shifts with MoE-LayerNorm
by: Fan, Xiao, et al.
Published: (2025)
by: Fan, Xiao, et al.
Published: (2025)
GeoNorm: Unify Pre-Norm and Post-Norm with Geodesic Optimization
by: Zheng, Chuanyang, et al.
Published: (2026)
by: Zheng, Chuanyang, et al.
Published: (2026)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
by: Wang, Wenxun, et al.
Published: (2025)
by: Wang, Wenxun, et al.
Published: (2025)
Does Machine Unlearning Truly Remove Knowledge?
by: Chen, Haokun, et al.
Published: (2025)
by: Chen, Haokun, et al.
Published: (2025)
Implicit Word Reordering with Knowledge Distillation for Cross-Lingual Dependency Parsing
by: Li, Zhuoran, et al.
Published: (2025)
by: Li, Zhuoran, et al.
Published: (2025)
Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization
by: Nrusimha, Aniruddha, et al.
Published: (2024)
by: Nrusimha, Aniruddha, et al.
Published: (2024)
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias
by: Huang, Ruiquan, et al.
Published: (2025)
by: Huang, Ruiquan, et al.
Published: (2025)
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
by: Skean, Oscar, et al.
Published: (2024)
by: Skean, Oscar, et al.
Published: (2024)
Impact of Layer Norm on Memorization and Generalization in Transformers
by: Singhal, Rishi, et al.
Published: (2025)
by: Singhal, Rishi, et al.
Published: (2025)
When and Why Does Unsupervised RL Succeed in Mathematical Reasoning? A Manifold Envelopment Perspective
by: Zhang, Zelin, et al.
Published: (2026)
by: Zhang, Zelin, et al.
Published: (2026)
When Does Divide and Conquer Work for Long Context LLM? A Noise Decomposition Framework
by: Xu, Zhen, et al.
Published: (2025)
by: Xu, Zhen, et al.
Published: (2025)
Does Visual Rendering Bypass Tokenization? Investigating Script-Tokenizer Misalignment in Pixel-Based Language Models
by: Susanto, Lucky, et al.
Published: (2026)
by: Susanto, Lucky, et al.
Published: (2026)
When Does Multimodality Lead to Better Time Series Forecasting?
by: Zhang, Xiyuan, et al.
Published: (2025)
by: Zhang, Xiyuan, et al.
Published: (2025)
When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars
by: Higuchi, Rei, et al.
Published: (2025)
by: Higuchi, Rei, et al.
Published: (2025)
SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration
by: Wen, Zhuofan, et al.
Published: (2026)
by: Wen, Zhuofan, et al.
Published: (2026)
Does Few-Shot Learning Help LLM Performance in Code Synthesis?
by: Xu, Derek, et al.
Published: (2024)
by: Xu, Derek, et al.
Published: (2024)
When is Tree Search Useful for LLM Planning? It Depends on the Discriminator
by: Chen, Ziru, et al.
Published: (2024)
by: Chen, Ziru, et al.
Published: (2024)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
by: Zheng, Mingqian, et al.
Published: (2023)
by: Zheng, Mingqian, et al.
Published: (2023)
On the Mathematical Relationship Between Layer Normalization and Dynamic Activation Functions
by: Stollenwerk, Felix
Published: (2025)
by: Stollenwerk, Felix
Published: (2025)
When Does a Language Model Commit? A Finite-Answer Theory of Pre-Verbalization Commitment
by: Zhang, Long, et al.
Published: (2026)
by: Zhang, Long, et al.
Published: (2026)
When Models Examine Themselves: Vocabulary-Activation Correspondence in Self-Referential Processing
by: Dadfar, Zachary Pedram
Published: (2026)
by: Dadfar, Zachary Pedram
Published: (2026)
HINT: Helping Ineffective Rollouts Navigate Towards Effectiveness
by: Wang, Xinyi, et al.
Published: (2025)
by: Wang, Xinyi, et al.
Published: (2025)
When Does Demographic Information Help? Data and Modeling Regimes for Perspective-Aware Hate Speech Detection
by: Cai, Weibin, et al.
Published: (2026)
by: Cai, Weibin, et al.
Published: (2026)
PolyNorm: Few-Shot LLM-Based Text Normalization for Text-to-Speech
by: Wong, Michel, et al.
Published: (2025)
by: Wong, Michel, et al.
Published: (2025)
Safety Training Persists Through Helpfulness Optimization in LLM Agents
by: Plaut, Benjamin
Published: (2026)
by: Plaut, Benjamin
Published: (2026)
Mimetic Initialization Helps State Space Models Learn to Recall
by: Trockman, Asher, et al.
Published: (2024)
by: Trockman, Asher, et al.
Published: (2024)
Skip-Connected Policy Optimization for Implicit Advantage
by: Teng, Fengwei, et al.
Published: (2026)
by: Teng, Fengwei, et al.
Published: (2026)
Bootstrapping Language Models with DPO Implicit Rewards
by: Chen, Changyu, et al.
Published: (2024)
by: Chen, Changyu, et al.
Published: (2024)
Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents
by: Talokar, Nivya, et al.
Published: (2026)
by: Talokar, Nivya, et al.
Published: (2026)
Similar Items
-
LayerNorm Induces Recency Bias in Transformer Decoders
by: Kim, Junu, et al.
Published: (2025) -
SLaNC: Static LayerNorm Calibration
by: Salmani, Mahsa, et al.
Published: (2024) -
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
by: Chen, Chen, et al.
Published: (2026) -
You can remove GPT2's LayerNorm by fine-tuning
by: Heimersheim, Stefan
Published: (2024) -
LayerNorm: A key component in parameter-efficient fine-tuning
by: ValizadehAslani, Taha, et al.
Published: (2024)