On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866913899628986368 |
|---|---|
| author | Diep, Nghiem T. Nguyen, Huy Nguyen, Chau Le, Minh Nguyen, Duy M. H. Sonntag, Daniel Niepert, Mathias Ho, Nhat |
| author_facet | Diep, Nghiem T. Nguyen, Huy Nguyen, Chau Le, Minh Nguyen, Duy M. H. Sonntag, Daniel Niepert, Mathias Ho, Nhat |
| contents | The LLaMA-Adapter has recently emerged as an efficient fine-tuning technique for LLaMA models, leveraging zero-initialized attention to stabilize training and enhance performance. However, despite its empirical success, the theoretical foundations of zero-initialized attention remain largely unexplored. In this paper, we provide a rigorous theoretical analysis, establishing a connection between zero-initialized attention and mixture-of-expert models. We prove that both linear and non-linear prompts, along with gating functions, can be optimally estimated, with non-linear prompts offering greater flexibility for future applications. Empirically, we validate our findings on the open LLM benchmarks, demonstrating that non-linear prompts outperform linear ones. Notably, even with limited training data, both prompt types consistently surpass vanilla attention, highlighting the robustness and adaptability of zero-initialized attention. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_03029 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation Diep, Nghiem T. Nguyen, Huy Nguyen, Chau Le, Minh Nguyen, Duy M. H. Sonntag, Daniel Niepert, Mathias Ho, Nhat Machine Learning The LLaMA-Adapter has recently emerged as an efficient fine-tuning technique for LLaMA models, leveraging zero-initialized attention to stabilize training and enhance performance. However, despite its empirical success, the theoretical foundations of zero-initialized attention remain largely unexplored. In this paper, we provide a rigorous theoretical analysis, establishing a connection between zero-initialized attention and mixture-of-expert models. We prove that both linear and non-linear prompts, along with gating functions, can be optimally estimated, with non-linear prompts offering greater flexibility for future applications. Empirically, we validate our findings on the open LLM benchmarks, demonstrating that non-linear prompts outperform linear ones. Notably, even with limited training data, both prompt types consistently surpass vanilla attention, highlighting the robustness and adaptability of zero-initialized attention. |
| title | On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2502.03029 |