On the Power of Context-Enhanced Learning in LLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhu, Xingyu, Panigrahi, Abhishek, Arora, Sanjeev
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909638189907968
author Zhu, Xingyu
Panigrahi, Abhishek
Arora, Sanjeev
author_facet Zhu, Xingyu
Panigrahi, Abhishek
Arora, Sanjeev
contents We formalize a new concept for LLMs, context-enhanced learning. It involves standard gradient-based learning on text except that the context is enhanced with additional data on which no auto-regressive gradients are computed. This setting is a gradient-based analog of usual in-context learning (ICL) and appears in some recent works. Using a multi-step reasoning task, we prove in a simplified setting that context-enhanced learning can be exponentially more sample-efficient than standard learning when the model is capable of ICL. At a mechanistic level, we find that the benefit of context-enhancement arises from a more accurate gradient learning signal. We also experimentally demonstrate that it appears hard to detect or recover learning materials that were used in the context during training. This may have implications for data security as well as copyright.
format Preprint
id arxiv_https___arxiv_org_abs_2503_01821
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On the Power of Context-Enhanced Learning in LLMs
Zhu, Xingyu
Panigrahi, Abhishek
Arora, Sanjeev
Machine Learning
We formalize a new concept for LLMs, context-enhanced learning. It involves standard gradient-based learning on text except that the context is enhanced with additional data on which no auto-regressive gradients are computed. This setting is a gradient-based analog of usual in-context learning (ICL) and appears in some recent works. Using a multi-step reasoning task, we prove in a simplified setting that context-enhanced learning can be exponentially more sample-efficient than standard learning when the model is capable of ICL. At a mechanistic level, we find that the benefit of context-enhancement arises from a more accurate gradient learning signal. We also experimentally demonstrate that it appears hard to detect or recover learning materials that were used in the context during training. This may have implications for data security as well as copyright.
title On the Power of Context-Enhanced Learning in LLMs
topic Machine Learning
url https://arxiv.org/abs/2503.01821