You can remove GPT2's LayerNorm by fine-tuning
Fuente:
arXiv
Saved in:
| Main Author: | Heimersheim, Stefan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LayerNorm: A key component in parameter-efficient fine-tuning
by: ValizadehAslani, Taha, et al.
Published: (2024)
by: ValizadehAslani, Taha, et al.
Published: (2024)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
by: Baroni, Luca, et al.
Published: (2025)
by: Baroni, Luca, et al.
Published: (2025)
LayerNorm Induces Recency Bias in Transformer Decoders
by: Kim, Junu, et al.
Published: (2025)
by: Kim, Junu, et al.
Published: (2025)
SLaNC: Static LayerNorm Calibration
by: Salmani, Mahsa, et al.
Published: (2024)
by: Salmani, Mahsa, et al.
Published: (2024)
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer
by: Verma, Lucky
Published: (2026)
by: Verma, Lucky
Published: (2026)
Geometry and Dynamics of LayerNorm
by: Riechers, Paul M.
Published: (2024)
by: Riechers, Paul M.
Published: (2024)
On the Role of Attention Masks and LayerNorm in Transformers
by: Wu, Xinyi, et al.
Published: (2024)
by: Wu, Xinyi, et al.
Published: (2024)
Does fine-tuning GPT-3 with the OpenAI API leak personally-identifiable information?
by: Sun, Albert Yu, et al.
Published: (2023)
by: Sun, Albert Yu, et al.
Published: (2023)
Complexity-aware fine-tuning
by: Goncharov, Andrey, et al.
Published: (2025)
by: Goncharov, Andrey, et al.
Published: (2025)
Replaying pre-training data improves fine-tuning
by: Kotha, Suhas, et al.
Published: (2026)
by: Kotha, Suhas, et al.
Published: (2026)
Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
by: Zhong, Ziqian, et al.
Published: (2025)
by: Zhong, Ziqian, et al.
Published: (2025)
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
by: Lee, Daniel J., et al.
Published: (2024)
by: Lee, Daniel J., et al.
Published: (2024)
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
by: Qian, Wenhao, et al.
Published: (2025)
by: Qian, Wenhao, et al.
Published: (2025)
MoETTA: Test-Time Adaptation Under Mixed Distribution Shifts with MoE-LayerNorm
by: Fan, Xiao, et al.
Published: (2025)
by: Fan, Xiao, et al.
Published: (2025)
The representation landscape of few-shot learning and fine-tuning in large language models
by: Doimo, Diego, et al.
Published: (2024)
by: Doimo, Diego, et al.
Published: (2024)
The more polypersonal the better -- a short look on space geometry of fine-tuned layers
by: Kudriashov, Sergei, et al.
Published: (2025)
by: Kudriashov, Sergei, et al.
Published: (2025)
Deep literature reviews: an application of fine-tuned language models to migration research
by: Iacus, Stefano M., et al.
Published: (2025)
by: Iacus, Stefano M., et al.
Published: (2025)
GeoNorm: Unify Pre-Norm and Post-Norm with Geodesic Optimization
by: Zheng, Chuanyang, et al.
Published: (2026)
by: Zheng, Chuanyang, et al.
Published: (2026)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
by: Wang, Wenxun, et al.
Published: (2025)
by: Wang, Wenxun, et al.
Published: (2025)
Rethinking harmless refusals when fine-tuning foundation models
by: Pop, Florin, et al.
Published: (2024)
by: Pop, Florin, et al.
Published: (2024)
Uncertainty quantification in fine-tuned LLMs using LoRA ensembles
by: Balabanov, Oleksandr, et al.
Published: (2024)
by: Balabanov, Oleksandr, et al.
Published: (2024)
Multi-task retriever fine-tuning for domain-specific and efficient RAG
by: Béchard, Patrice, et al.
Published: (2025)
by: Béchard, Patrice, et al.
Published: (2025)
Improving embedding with contrastive fine-tuning on small datasets with expert-augmented scores
by: Lu, Jun, et al.
Published: (2024)
by: Lu, Jun, et al.
Published: (2024)
Impact of Layer Norm on Memorization and Generalization in Transformers
by: Singhal, Rishi, et al.
Published: (2025)
by: Singhal, Rishi, et al.
Published: (2025)
Fine-tuning MLLMs Without Forgetting Is Easier Than You Think
by: Li, He, et al.
Published: (2026)
by: Li, He, et al.
Published: (2026)
Fine-tuning can Help Detect Pretraining Data from Large Language Models
by: Zhang, Hengxiang, et al.
Published: (2024)
by: Zhang, Hengxiang, et al.
Published: (2024)
What explains the success of cross-modal fine-tuning with ORCA?
by: García-de-Herreros, Paloma, et al.
Published: (2024)
by: García-de-Herreros, Paloma, et al.
Published: (2024)
Layer-wise Importance Matters: Less Memory for Better Performance in Parameter-efficient Fine-tuning of Large Language Models
by: Yao, Kai, et al.
Published: (2024)
by: Yao, Kai, et al.
Published: (2024)
Investigating the performance of Retrieval-Augmented Generation and fine-tuning for the development of AI-driven knowledge-based systems
by: Lakatos, Robert, et al.
Published: (2024)
by: Lakatos, Robert, et al.
Published: (2024)
Enhancing Q&A Text Retrieval with Ranking Models: Benchmarking, fine-tuning and deploying Rerankers for RAG
by: Moreira, Gabriel de Souza P., et al.
Published: (2024)
by: Moreira, Gabriel de Souza P., et al.
Published: (2024)
GRASS: Gradient-based Adaptive Layer-wise Importance Sampling for Memory-efficient Large Language Model Fine-tuning
by: Tian, Kaiyuan, et al.
Published: (2026)
by: Tian, Kaiyuan, et al.
Published: (2026)
JMI at SemEval 2024 Task 3: Two-step approach for multimodal ECAC using in-context learning with GPT and instruction-tuned Llama models
by: Arefa, et al.
Published: (2024)
by: Arefa, et al.
Published: (2024)
MaxPoolBERT: Enhancing BERT Classification via Layer- and Token-Wise Aggregation
by: Behrendt, Maike, et al.
Published: (2025)
by: Behrendt, Maike, et al.
Published: (2025)
LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report
by: Zhao, Justin, et al.
Published: (2024)
by: Zhao, Justin, et al.
Published: (2024)
Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
by: Zhou, Huichi, et al.
Published: (2025)
by: Zhou, Huichi, et al.
Published: (2025)
MiniGPT: Rebuilding GPT from First Principles
by: Joseph, Jibin
Published: (2026)
by: Joseph, Jibin
Published: (2026)
TransactionGPT
by: Dou, Yingtong, et al.
Published: (2025)
by: Dou, Yingtong, et al.
Published: (2025)
ZzzGPT: An Interactive GPT Approach to Enhance Sleep Quality
by: Khaokaew, Yonchanok, et al.
Published: (2023)
by: Khaokaew, Yonchanok, et al.
Published: (2023)
Evaluating Synthetic Activations composed of SAE Latents in GPT-2
by: Giglemiani, Giorgi, et al.
Published: (2024)
by: Giglemiani, Giorgi, et al.
Published: (2024)
Similar Items
-
LayerNorm: A key component in parameter-efficient fine-tuning
by: ValizadehAslani, Taha, et al.
Published: (2024) -
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
by: Baroni, Luca, et al.
Published: (2025) -
LayerNorm Induces Recency Bias in Transformer Decoders
by: Kim, Junu, et al.
Published: (2025) -
SLaNC: Static LayerNorm Calibration
by: Salmani, Mahsa, et al.
Published: (2024) -
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
by: Chen, Chen, et al.
Published: (2026)