Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiong, Zheyang, Cai, Ziyang, Cooper, John, Ge, Albert, Papageorgiou, Vasilis, Sifakis, Zack, Giannou, Angeliki, Lin, Ziqian, Yang, Liu, Agarwal, Saurabh, Chrysos, Grigorios G, Oymak, Samet, Lee, Kangwook, Papailiopoulos, Dimitris
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909340104916992
author Xiong, Zheyang
Cai, Ziyang
Cooper, John
Ge, Albert
Papageorgiou, Vasilis
Sifakis, Zack
Giannou, Angeliki
Lin, Ziqian
Yang, Liu
Agarwal, Saurabh
Chrysos, Grigorios G
Oymak, Samet
Lee, Kangwook
Papailiopoulos, Dimitris
author_facet Xiong, Zheyang
Cai, Ziyang
Cooper, John
Ge, Albert
Papageorgiou, Vasilis
Sifakis, Zack
Giannou, Angeliki
Lin, Ziqian
Yang, Liu
Agarwal, Saurabh
Chrysos, Grigorios G
Oymak, Samet
Lee, Kangwook
Papailiopoulos, Dimitris
contents Large Language Models (LLMs) have demonstrated remarkable in-context learning (ICL) capabilities. In this study, we explore a surprising phenomenon related to ICL: LLMs can perform multiple, computationally distinct ICL tasks simultaneously, during a single inference call, a capability we term "task superposition". We provide empirical evidence of this phenomenon across various LLM families and scales and show that this phenomenon emerges even if we train the model to in-context learn one task at a time. We offer theoretical explanations that this capability is well within the expressive power of transformers. We also explore how LLMs internally compose task vectors during superposition. Furthermore, we show that larger models can solve more ICL tasks in parallel, and better calibrate their output distribution. Our findings offer insights into the latent capabilities of LLMs, further substantiate the perspective of "LLMs as superposition of simulators", and raise questions about the mechanisms enabling simultaneous task execution.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05603
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition
Xiong, Zheyang
Cai, Ziyang
Cooper, John
Ge, Albert
Papageorgiou, Vasilis
Sifakis, Zack
Giannou, Angeliki
Lin, Ziqian
Yang, Liu
Agarwal, Saurabh
Chrysos, Grigorios G
Oymak, Samet
Lee, Kangwook
Papailiopoulos, Dimitris
Machine Learning
Artificial Intelligence
Computation and Language
Large Language Models (LLMs) have demonstrated remarkable in-context learning (ICL) capabilities. In this study, we explore a surprising phenomenon related to ICL: LLMs can perform multiple, computationally distinct ICL tasks simultaneously, during a single inference call, a capability we term "task superposition". We provide empirical evidence of this phenomenon across various LLM families and scales and show that this phenomenon emerges even if we train the model to in-context learn one task at a time. We offer theoretical explanations that this capability is well within the expressive power of transformers. We also explore how LLMs internally compose task vectors during superposition. Furthermore, we show that larger models can solve more ICL tasks in parallel, and better calibrate their output distribution. Our findings offer insights into the latent capabilities of LLMs, further substantiate the perspective of "LLMs as superposition of simulators", and raise questions about the mechanisms enabling simultaneous task execution.
title Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.05603