You can remove GPT2's LayerNorm by fine-tuning
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Heimersheim, Stefan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LayerNorm: A key component in parameter-efficient fine-tuning
von: ValizadehAslani, Taha, et al.
Veröffentlicht: (2024)
von: ValizadehAslani, Taha, et al.
Veröffentlicht: (2024)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
von: Baroni, Luca, et al.
Veröffentlicht: (2025)
von: Baroni, Luca, et al.
Veröffentlicht: (2025)
LayerNorm Induces Recency Bias in Transformer Decoders
von: Kim, Junu, et al.
Veröffentlicht: (2025)
von: Kim, Junu, et al.
Veröffentlicht: (2025)
SLaNC: Static LayerNorm Calibration
von: Salmani, Mahsa, et al.
Veröffentlicht: (2024)
von: Salmani, Mahsa, et al.
Veröffentlicht: (2024)
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
von: Chen, Chen, et al.
Veröffentlicht: (2026)
von: Chen, Chen, et al.
Veröffentlicht: (2026)
When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer
von: Verma, Lucky
Veröffentlicht: (2026)
von: Verma, Lucky
Veröffentlicht: (2026)
Geometry and Dynamics of LayerNorm
von: Riechers, Paul M.
Veröffentlicht: (2024)
von: Riechers, Paul M.
Veröffentlicht: (2024)
On the Role of Attention Masks and LayerNorm in Transformers
von: Wu, Xinyi, et al.
Veröffentlicht: (2024)
von: Wu, Xinyi, et al.
Veröffentlicht: (2024)
Does fine-tuning GPT-3 with the OpenAI API leak personally-identifiable information?
von: Sun, Albert Yu, et al.
Veröffentlicht: (2023)
von: Sun, Albert Yu, et al.
Veröffentlicht: (2023)
Complexity-aware fine-tuning
von: Goncharov, Andrey, et al.
Veröffentlicht: (2025)
von: Goncharov, Andrey, et al.
Veröffentlicht: (2025)
Replaying pre-training data improves fine-tuning
von: Kotha, Suhas, et al.
Veröffentlicht: (2026)
von: Kotha, Suhas, et al.
Veröffentlicht: (2026)
Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
von: Zhong, Ziqian, et al.
Veröffentlicht: (2025)
von: Zhong, Ziqian, et al.
Veröffentlicht: (2025)
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
von: Lee, Daniel J., et al.
Veröffentlicht: (2024)
von: Lee, Daniel J., et al.
Veröffentlicht: (2024)
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
von: Qian, Wenhao, et al.
Veröffentlicht: (2025)
von: Qian, Wenhao, et al.
Veröffentlicht: (2025)
MoETTA: Test-Time Adaptation Under Mixed Distribution Shifts with MoE-LayerNorm
von: Fan, Xiao, et al.
Veröffentlicht: (2025)
von: Fan, Xiao, et al.
Veröffentlicht: (2025)
The representation landscape of few-shot learning and fine-tuning in large language models
von: Doimo, Diego, et al.
Veröffentlicht: (2024)
von: Doimo, Diego, et al.
Veröffentlicht: (2024)
The more polypersonal the better -- a short look on space geometry of fine-tuned layers
von: Kudriashov, Sergei, et al.
Veröffentlicht: (2025)
von: Kudriashov, Sergei, et al.
Veröffentlicht: (2025)
Deep literature reviews: an application of fine-tuned language models to migration research
von: Iacus, Stefano M., et al.
Veröffentlicht: (2025)
von: Iacus, Stefano M., et al.
Veröffentlicht: (2025)
GeoNorm: Unify Pre-Norm and Post-Norm with Geodesic Optimization
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2026)
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2026)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
von: Wang, Wenxun, et al.
Veröffentlicht: (2025)
von: Wang, Wenxun, et al.
Veröffentlicht: (2025)
Rethinking harmless refusals when fine-tuning foundation models
von: Pop, Florin, et al.
Veröffentlicht: (2024)
von: Pop, Florin, et al.
Veröffentlicht: (2024)
Uncertainty quantification in fine-tuned LLMs using LoRA ensembles
von: Balabanov, Oleksandr, et al.
Veröffentlicht: (2024)
von: Balabanov, Oleksandr, et al.
Veröffentlicht: (2024)
Multi-task retriever fine-tuning for domain-specific and efficient RAG
von: Béchard, Patrice, et al.
Veröffentlicht: (2025)
von: Béchard, Patrice, et al.
Veröffentlicht: (2025)
Improving embedding with contrastive fine-tuning on small datasets with expert-augmented scores
von: Lu, Jun, et al.
Veröffentlicht: (2024)
von: Lu, Jun, et al.
Veröffentlicht: (2024)
Impact of Layer Norm on Memorization and Generalization in Transformers
von: Singhal, Rishi, et al.
Veröffentlicht: (2025)
von: Singhal, Rishi, et al.
Veröffentlicht: (2025)
Fine-tuning MLLMs Without Forgetting Is Easier Than You Think
von: Li, He, et al.
Veröffentlicht: (2026)
von: Li, He, et al.
Veröffentlicht: (2026)
Fine-tuning can Help Detect Pretraining Data from Large Language Models
von: Zhang, Hengxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Hengxiang, et al.
Veröffentlicht: (2024)
What explains the success of cross-modal fine-tuning with ORCA?
von: García-de-Herreros, Paloma, et al.
Veröffentlicht: (2024)
von: García-de-Herreros, Paloma, et al.
Veröffentlicht: (2024)
Layer-wise Importance Matters: Less Memory for Better Performance in Parameter-efficient Fine-tuning of Large Language Models
von: Yao, Kai, et al.
Veröffentlicht: (2024)
von: Yao, Kai, et al.
Veröffentlicht: (2024)
Investigating the performance of Retrieval-Augmented Generation and fine-tuning for the development of AI-driven knowledge-based systems
von: Lakatos, Robert, et al.
Veröffentlicht: (2024)
von: Lakatos, Robert, et al.
Veröffentlicht: (2024)
Enhancing Q&A Text Retrieval with Ranking Models: Benchmarking, fine-tuning and deploying Rerankers for RAG
von: Moreira, Gabriel de Souza P., et al.
Veröffentlicht: (2024)
von: Moreira, Gabriel de Souza P., et al.
Veröffentlicht: (2024)
GRASS: Gradient-based Adaptive Layer-wise Importance Sampling for Memory-efficient Large Language Model Fine-tuning
von: Tian, Kaiyuan, et al.
Veröffentlicht: (2026)
von: Tian, Kaiyuan, et al.
Veröffentlicht: (2026)
JMI at SemEval 2024 Task 3: Two-step approach for multimodal ECAC using in-context learning with GPT and instruction-tuned Llama models
von: Arefa, et al.
Veröffentlicht: (2024)
von: Arefa, et al.
Veröffentlicht: (2024)
MaxPoolBERT: Enhancing BERT Classification via Layer- and Token-Wise Aggregation
von: Behrendt, Maike, et al.
Veröffentlicht: (2025)
von: Behrendt, Maike, et al.
Veröffentlicht: (2025)
LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report
von: Zhao, Justin, et al.
Veröffentlicht: (2024)
von: Zhao, Justin, et al.
Veröffentlicht: (2024)
Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
von: Zhou, Huichi, et al.
Veröffentlicht: (2025)
von: Zhou, Huichi, et al.
Veröffentlicht: (2025)
MiniGPT: Rebuilding GPT from First Principles
von: Joseph, Jibin
Veröffentlicht: (2026)
von: Joseph, Jibin
Veröffentlicht: (2026)
TransactionGPT
von: Dou, Yingtong, et al.
Veröffentlicht: (2025)
von: Dou, Yingtong, et al.
Veröffentlicht: (2025)
ZzzGPT: An Interactive GPT Approach to Enhance Sleep Quality
von: Khaokaew, Yonchanok, et al.
Veröffentlicht: (2023)
von: Khaokaew, Yonchanok, et al.
Veröffentlicht: (2023)
Evaluating Synthetic Activations composed of SAE Latents in GPT-2
von: Giglemiani, Giorgi, et al.
Veröffentlicht: (2024)
von: Giglemiani, Giorgi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LayerNorm: A key component in parameter-efficient fine-tuning
von: ValizadehAslani, Taha, et al.
Veröffentlicht: (2024) -
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
von: Baroni, Luca, et al.
Veröffentlicht: (2025) -
LayerNorm Induces Recency Bias in Transformer Decoders
von: Kim, Junu, et al.
Veröffentlicht: (2025) -
SLaNC: Static LayerNorm Calibration
von: Salmani, Mahsa, et al.
Veröffentlicht: (2024) -
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
von: Chen, Chen, et al.
Veröffentlicht: (2026)