In-context learning and Occam's razor

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Elmoznino, Eric, Marty, Tom, Kasetty, Tejas, Gagnon, Leo, Mittal, Sarthak, Fathi, Mahan, Sridhar, Dhanya, Lajoie, Guillaume
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912409989414912
author Elmoznino, Eric
Marty, Tom
Kasetty, Tejas
Gagnon, Leo
Mittal, Sarthak
Fathi, Mahan
Sridhar, Dhanya
Lajoie, Guillaume
author_facet Elmoznino, Eric
Marty, Tom
Kasetty, Tejas
Gagnon, Leo
Mittal, Sarthak
Fathi, Mahan
Sridhar, Dhanya
Lajoie, Guillaume
contents A central goal of machine learning is generalization. While the No Free Lunch Theorem states that we cannot obtain theoretical guarantees for generalization without further assumptions, in practice we observe that simple models which explain the training data generalize best: a principle called Occam's razor. Despite the need for simple models, most current approaches in machine learning only minimize the training error, and at best indirectly promote simplicity through regularization or architecture design. Here, we draw a connection between Occam's razor and in-context learning: an emergent ability of certain sequence models like Transformers to learn at inference time from past observations in a sequence. In particular, we show that the next-token prediction loss used to train in-context learners is directly equivalent to a data compression technique called prequential coding, and that minimizing this loss amounts to jointly minimizing both the training error and the complexity of the model that was implicitly learned from context. Our theory and the empirical experiments we use to support it not only provide a normative account of in-context learning, but also elucidate the shortcomings of current in-context learning methods, suggesting ways in which they can be improved. We make our code available at https://github.com/3rdCore/PrequentialCode.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14086
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle In-context learning and Occam's razor
Elmoznino, Eric
Marty, Tom
Kasetty, Tejas
Gagnon, Leo
Mittal, Sarthak
Fathi, Mahan
Sridhar, Dhanya
Lajoie, Guillaume
Machine Learning
Artificial Intelligence
A central goal of machine learning is generalization. While the No Free Lunch Theorem states that we cannot obtain theoretical guarantees for generalization without further assumptions, in practice we observe that simple models which explain the training data generalize best: a principle called Occam's razor. Despite the need for simple models, most current approaches in machine learning only minimize the training error, and at best indirectly promote simplicity through regularization or architecture design. Here, we draw a connection between Occam's razor and in-context learning: an emergent ability of certain sequence models like Transformers to learn at inference time from past observations in a sequence. In particular, we show that the next-token prediction loss used to train in-context learners is directly equivalent to a data compression technique called prequential coding, and that minimizing this loss amounts to jointly minimizing both the training error and the complexity of the model that was implicitly learned from context. Our theory and the empirical experiments we use to support it not only provide a normative account of in-context learning, but also elucidate the shortcomings of current in-context learning methods, suggesting ways in which they can be improved. We make our code available at https://github.com/3rdCore/PrequentialCode.
title In-context learning and Occam's razor
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.14086