LayerNorm: A key component in parameter-efficient fine-tuning
Fuente:
arXiv
Salvato in:
| Autori principali: | ValizadehAslani, Taha, Liang, Hualou |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
You can remove GPT2's LayerNorm by fine-tuning
di: Heimersheim, Stefan
Pubblicazione: (2024)
di: Heimersheim, Stefan
Pubblicazione: (2024)
LayerNorm Induces Recency Bias in Transformer Decoders
di: Kim, Junu, et al.
Pubblicazione: (2025)
di: Kim, Junu, et al.
Pubblicazione: (2025)
SLaNC: Static LayerNorm Calibration
di: Salmani, Mahsa, et al.
Pubblicazione: (2024)
di: Salmani, Mahsa, et al.
Pubblicazione: (2024)
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
di: Chen, Chen, et al.
Pubblicazione: (2026)
di: Chen, Chen, et al.
Pubblicazione: (2026)
When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer
di: Verma, Lucky
Pubblicazione: (2026)
di: Verma, Lucky
Pubblicazione: (2026)
Geometry and Dynamics of LayerNorm
di: Riechers, Paul M.
Pubblicazione: (2024)
di: Riechers, Paul M.
Pubblicazione: (2024)
On the Role of Attention Masks and LayerNorm in Transformers
di: Wu, Xinyi, et al.
Pubblicazione: (2024)
di: Wu, Xinyi, et al.
Pubblicazione: (2024)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
di: Baroni, Luca, et al.
Pubblicazione: (2025)
di: Baroni, Luca, et al.
Pubblicazione: (2025)
Replaying pre-training data improves fine-tuning
di: Kotha, Suhas, et al.
Pubblicazione: (2026)
di: Kotha, Suhas, et al.
Pubblicazione: (2026)
Complexity-aware fine-tuning
di: Goncharov, Andrey, et al.
Pubblicazione: (2025)
di: Goncharov, Andrey, et al.
Pubblicazione: (2025)
Multi-task retriever fine-tuning for domain-specific and efficient RAG
di: Béchard, Patrice, et al.
Pubblicazione: (2025)
di: Béchard, Patrice, et al.
Pubblicazione: (2025)
Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
di: Zhong, Ziqian, et al.
Pubblicazione: (2025)
di: Zhong, Ziqian, et al.
Pubblicazione: (2025)
Layer-wise Importance Matters: Less Memory for Better Performance in Parameter-efficient Fine-tuning of Large Language Models
di: Yao, Kai, et al.
Pubblicazione: (2024)
di: Yao, Kai, et al.
Pubblicazione: (2024)
GRASS: Gradient-based Adaptive Layer-wise Importance Sampling for Memory-efficient Large Language Model Fine-tuning
di: Tian, Kaiyuan, et al.
Pubblicazione: (2026)
di: Tian, Kaiyuan, et al.
Pubblicazione: (2026)
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
di: Qian, Wenhao, et al.
Pubblicazione: (2025)
di: Qian, Wenhao, et al.
Pubblicazione: (2025)
MoETTA: Test-Time Adaptation Under Mixed Distribution Shifts with MoE-LayerNorm
di: Fan, Xiao, et al.
Pubblicazione: (2025)
di: Fan, Xiao, et al.
Pubblicazione: (2025)
The representation landscape of few-shot learning and fine-tuning in large language models
di: Doimo, Diego, et al.
Pubblicazione: (2024)
di: Doimo, Diego, et al.
Pubblicazione: (2024)
The more polypersonal the better -- a short look on space geometry of fine-tuned layers
di: Kudriashov, Sergei, et al.
Pubblicazione: (2025)
di: Kudriashov, Sergei, et al.
Pubblicazione: (2025)
Deep literature reviews: an application of fine-tuned language models to migration research
di: Iacus, Stefano M., et al.
Pubblicazione: (2025)
di: Iacus, Stefano M., et al.
Pubblicazione: (2025)
Does fine-tuning GPT-3 with the OpenAI API leak personally-identifiable information?
di: Sun, Albert Yu, et al.
Pubblicazione: (2023)
di: Sun, Albert Yu, et al.
Pubblicazione: (2023)
GeoNorm: Unify Pre-Norm and Post-Norm with Geodesic Optimization
di: Zheng, Chuanyang, et al.
Pubblicazione: (2026)
di: Zheng, Chuanyang, et al.
Pubblicazione: (2026)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
di: Wang, Wenxun, et al.
Pubblicazione: (2025)
di: Wang, Wenxun, et al.
Pubblicazione: (2025)
Rethinking harmless refusals when fine-tuning foundation models
di: Pop, Florin, et al.
Pubblicazione: (2024)
di: Pop, Florin, et al.
Pubblicazione: (2024)
AtP*: An efficient and scalable method for localizing LLM behaviour to components
di: Kramár, János, et al.
Pubblicazione: (2024)
di: Kramár, János, et al.
Pubblicazione: (2024)
Uncertainty quantification in fine-tuned LLMs using LoRA ensembles
di: Balabanov, Oleksandr, et al.
Pubblicazione: (2024)
di: Balabanov, Oleksandr, et al.
Pubblicazione: (2024)
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
di: Ben-Zaken, Elad, et al.
Pubblicazione: (2021)
di: Ben-Zaken, Elad, et al.
Pubblicazione: (2021)
Improving embedding with contrastive fine-tuning on small datasets with expert-augmented scores
di: Lu, Jun, et al.
Pubblicazione: (2024)
di: Lu, Jun, et al.
Pubblicazione: (2024)
Quantum-PEFT: Ultra parameter-efficient fine-tuning
di: Koike-Akino, Toshiaki, et al.
Pubblicazione: (2025)
di: Koike-Akino, Toshiaki, et al.
Pubblicazione: (2025)
Impact of Layer Norm on Memorization and Generalization in Transformers
di: Singhal, Rishi, et al.
Pubblicazione: (2025)
di: Singhal, Rishi, et al.
Pubblicazione: (2025)
Enhancing Q&A Text Retrieval with Ranking Models: Benchmarking, fine-tuning and deploying Rerankers for RAG
di: Moreira, Gabriel de Souza P., et al.
Pubblicazione: (2024)
di: Moreira, Gabriel de Souza P., et al.
Pubblicazione: (2024)
Correct and Optimal: the Regular Expression Inference Challenge
di: Valizadeh, Mojtaba, et al.
Pubblicazione: (2023)
di: Valizadeh, Mojtaba, et al.
Pubblicazione: (2023)
What explains the success of cross-modal fine-tuning with ORCA?
di: García-de-Herreros, Paloma, et al.
Pubblicazione: (2024)
di: García-de-Herreros, Paloma, et al.
Pubblicazione: (2024)
Investigating the performance of Retrieval-Augmented Generation and fine-tuning for the development of AI-driven knowledge-based systems
di: Lakatos, Robert, et al.
Pubblicazione: (2024)
di: Lakatos, Robert, et al.
Pubblicazione: (2024)
Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
di: Zhou, Huichi, et al.
Pubblicazione: (2025)
di: Zhou, Huichi, et al.
Pubblicazione: (2025)
Fine-tuning Large Language Models for Domain-specific Machine Translation
di: Zheng, Jiawei, et al.
Pubblicazione: (2024)
di: Zheng, Jiawei, et al.
Pubblicazione: (2024)
Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers
di: Wong, Liang Ze
Pubblicazione: (2025)
di: Wong, Liang Ze
Pubblicazione: (2025)
Harmony in Divergence: Towards Fast, Accurate, and Memory-efficient Zeroth-order LLM Fine-tuning
di: Tan, Qitao, et al.
Pubblicazione: (2025)
di: Tan, Qitao, et al.
Pubblicazione: (2025)
NorMuon: Making Muon more efficient and scalable
di: Li, Zichong, et al.
Pubblicazione: (2025)
di: Li, Zichong, et al.
Pubblicazione: (2025)
Towards Higher Effective Rank in Parameter-efficient Fine-tuning using Khatri--Rao Product
di: Albert, Paul, et al.
Pubblicazione: (2025)
di: Albert, Paul, et al.
Pubblicazione: (2025)
Tied-Lora: Enhancing parameter efficiency of LoRA with weight tying
di: Renduchintala, Adithya, et al.
Pubblicazione: (2023)
di: Renduchintala, Adithya, et al.
Pubblicazione: (2023)
Documenti analoghi
-
You can remove GPT2's LayerNorm by fine-tuning
di: Heimersheim, Stefan
Pubblicazione: (2024) -
LayerNorm Induces Recency Bias in Transformer Decoders
di: Kim, Junu, et al.
Pubblicazione: (2025) -
SLaNC: Static LayerNorm Calibration
di: Salmani, Mahsa, et al.
Pubblicazione: (2024) -
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
di: Chen, Chen, et al.
Pubblicazione: (2026) -
When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer
di: Verma, Lucky
Pubblicazione: (2026)