On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Diep, Nghiem T., Nguyen, Huy, Nguyen, Chau, Le, Minh, Nguyen, Duy M. H., Sonntag, Daniel, Niepert, Mathias, Ho, Nhat
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913899628986368
author Diep, Nghiem T.
Nguyen, Huy
Nguyen, Chau
Le, Minh
Nguyen, Duy M. H.
Sonntag, Daniel
Niepert, Mathias
Ho, Nhat
author_facet Diep, Nghiem T.
Nguyen, Huy
Nguyen, Chau
Le, Minh
Nguyen, Duy M. H.
Sonntag, Daniel
Niepert, Mathias
Ho, Nhat
contents The LLaMA-Adapter has recently emerged as an efficient fine-tuning technique for LLaMA models, leveraging zero-initialized attention to stabilize training and enhance performance. However, despite its empirical success, the theoretical foundations of zero-initialized attention remain largely unexplored. In this paper, we provide a rigorous theoretical analysis, establishing a connection between zero-initialized attention and mixture-of-expert models. We prove that both linear and non-linear prompts, along with gating functions, can be optimally estimated, with non-linear prompts offering greater flexibility for future applications. Empirically, we validate our findings on the open LLM benchmarks, demonstrating that non-linear prompts outperform linear ones. Notably, even with limited training data, both prompt types consistently surpass vanilla attention, highlighting the robustness and adaptability of zero-initialized attention.
format Preprint
id arxiv_https___arxiv_org_abs_2502_03029
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation
Diep, Nghiem T.
Nguyen, Huy
Nguyen, Chau
Le, Minh
Nguyen, Duy M. H.
Sonntag, Daniel
Niepert, Mathias
Ho, Nhat
Machine Learning
The LLaMA-Adapter has recently emerged as an efficient fine-tuning technique for LLaMA models, leveraging zero-initialized attention to stabilize training and enhance performance. However, despite its empirical success, the theoretical foundations of zero-initialized attention remain largely unexplored. In this paper, we provide a rigorous theoretical analysis, establishing a connection between zero-initialized attention and mixture-of-expert models. We prove that both linear and non-linear prompts, along with gating functions, can be optimally estimated, with non-linear prompts offering greater flexibility for future applications. Empirically, we validate our findings on the open LLM benchmarks, demonstrating that non-linear prompts outperform linear ones. Notably, even with limited training data, both prompt types consistently surpass vanilla attention, highlighting the robustness and adaptability of zero-initialized attention.
title On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation
topic Machine Learning
url https://arxiv.org/abs/2502.03029