In-Context Learning with Long-Context Models: An In-Depth Exploration

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Bertsch, Amanda, Ivgi, Maor, Xiao, Emily, Alon, Uri, Berant, Jonathan, Gormley, Matthew R., Neubig, Graham
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929739835375616
author Bertsch, Amanda
Ivgi, Maor
Xiao, Emily
Alon, Uri
Berant, Jonathan
Gormley, Matthew R.
Neubig, Graham
author_facet Bertsch, Amanda
Ivgi, Maor
Xiao, Emily
Alon, Uri
Berant, Jonathan
Gormley, Matthew R.
Neubig, Graham
contents As model context lengths continue to increase, the number of demonstrations that can be provided in-context approaches the size of entire training datasets. We study the behavior of in-context learning (ICL) at this extreme scale on multiple datasets and models. We show that, for many datasets with large label spaces, performance continues to increase with thousands of demonstrations. We contrast this with example retrieval and finetuning: example retrieval shows excellent performance at low context lengths but has diminished gains with more demonstrations; finetuning is more data hungry than ICL but can exceed long-context ICL performance with additional data. We use the ICL setting to study several properties of both in-context learning and long-context models. We show that long-context ICL is less sensitive to random input shuffling than short-context ICL, that grouping of same-label examples negatively impacts performance, and that the performance boosts do not arise from cumulative gain from encoding many examples together. We conclude that long-context ICL can be an effective tool, and may not require long-context for encoding the demonstration set at all.
format Preprint
id arxiv_https___arxiv_org_abs_2405_00200
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle In-Context Learning with Long-Context Models: An In-Depth Exploration
Bertsch, Amanda
Ivgi, Maor
Xiao, Emily
Alon, Uri
Berant, Jonathan
Gormley, Matthew R.
Neubig, Graham
Computation and Language
As model context lengths continue to increase, the number of demonstrations that can be provided in-context approaches the size of entire training datasets. We study the behavior of in-context learning (ICL) at this extreme scale on multiple datasets and models. We show that, for many datasets with large label spaces, performance continues to increase with thousands of demonstrations. We contrast this with example retrieval and finetuning: example retrieval shows excellent performance at low context lengths but has diminished gains with more demonstrations; finetuning is more data hungry than ICL but can exceed long-context ICL performance with additional data. We use the ICL setting to study several properties of both in-context learning and long-context models. We show that long-context ICL is less sensitive to random input shuffling than short-context ICL, that grouping of same-label examples negatively impacts performance, and that the performance boosts do not arise from cumulative gain from encoding many examples together. We conclude that long-context ICL can be an effective tool, and may not require long-context for encoding the demonstration set at all.
title In-Context Learning with Long-Context Models: An In-Depth Exploration
topic Computation and Language
url https://arxiv.org/abs/2405.00200