Is attention required for ICL? Exploring the Relationship Between Model Architecture and In-Context Learning Ability

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lee, Ivan, Jiang, Nan, Berg-Kirkpatrick, Taylor
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911822445019136
author Lee, Ivan
Jiang, Nan
Berg-Kirkpatrick, Taylor
author_facet Lee, Ivan
Jiang, Nan
Berg-Kirkpatrick, Taylor
contents What is the relationship between model architecture and the ability to perform in-context learning? In this empirical study, we take the first steps toward answering this question. We evaluate thirteen model architectures capable of causal language modeling across a suite of synthetic in-context learning tasks. These selected architectures represent a broad range of paradigms, including recurrent and convolution-based neural networks, transformers, state space model inspired, and other emerging attention alternatives. We discover that all the considered architectures can perform in-context learning under a wider range of conditions than previously documented. Additionally, we observe stark differences in statistical efficiency and consistency by varying the number of in-context examples and task difficulty. We also measure each architecture's predisposition towards in-context learning when presented with the option to memorize rather than leverage in-context examples. Finally, and somewhat surprisingly, we find that several attention alternatives are sometimes competitive with or better in-context learners than transformers. However, no single architecture demonstrates consistency across all tasks, with performance either plateauing or declining when confronted with a significantly larger number of in-context examples than those encountered during gradient-based training.
format Preprint
id arxiv_https___arxiv_org_abs_2310_08049
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Is attention required for ICL? Exploring the Relationship Between Model Architecture and In-Context Learning Ability
Lee, Ivan
Jiang, Nan
Berg-Kirkpatrick, Taylor
Machine Learning
What is the relationship between model architecture and the ability to perform in-context learning? In this empirical study, we take the first steps toward answering this question. We evaluate thirteen model architectures capable of causal language modeling across a suite of synthetic in-context learning tasks. These selected architectures represent a broad range of paradigms, including recurrent and convolution-based neural networks, transformers, state space model inspired, and other emerging attention alternatives. We discover that all the considered architectures can perform in-context learning under a wider range of conditions than previously documented. Additionally, we observe stark differences in statistical efficiency and consistency by varying the number of in-context examples and task difficulty. We also measure each architecture's predisposition towards in-context learning when presented with the option to memorize rather than leverage in-context examples. Finally, and somewhat surprisingly, we find that several attention alternatives are sometimes competitive with or better in-context learners than transformers. However, no single architecture demonstrates consistency across all tasks, with performance either plateauing or declining when confronted with a significantly larger number of in-context examples than those encountered during gradient-based training.
title Is attention required for ICL? Exploring the Relationship Between Model Architecture and In-Context Learning Ability
topic Machine Learning
url https://arxiv.org/abs/2310.08049