Learning Task Representations from In-Context Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Saglam, Baturay, Hu, Xinyang, Yang, Zhuoran, Kalogerias, Dionysis, Karbasi, Amin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911254838247424
author Saglam, Baturay
Hu, Xinyang
Yang, Zhuoran
Kalogerias, Dionysis
Karbasi, Amin
author_facet Saglam, Baturay
Hu, Xinyang
Yang, Zhuoran
Kalogerias, Dionysis
Karbasi, Amin
contents Large language models (LLMs) have demonstrated remarkable proficiency in in-context learning (ICL), where models adapt to new tasks through example-based prompts without requiring parameter updates. However, understanding how tasks are internally encoded and generalized remains a challenge. To address some of the empirical and technical gaps in the literature, we introduce an automated formulation for encoding task information in ICL prompts as a function of attention heads within the transformer architecture. This approach computes a single task vector as a weighted sum of attention heads, with the weights optimized causally via gradient descent. Our findings show that existing methods fail to generalize effectively to modalities beyond text. In response, we also design a benchmark to evaluate whether a task vector can preserve task fidelity in functional regression tasks. The proposed method successfully extracts task-specific information from in-context demonstrations and excels in both text and regression tasks, demonstrating its generalizability across modalities.
format Preprint
id arxiv_https___arxiv_org_abs_2502_05390
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Task Representations from In-Context Learning
Saglam, Baturay
Hu, Xinyang
Yang, Zhuoran
Kalogerias, Dionysis
Karbasi, Amin
Computation and Language
Machine Learning
Large language models (LLMs) have demonstrated remarkable proficiency in in-context learning (ICL), where models adapt to new tasks through example-based prompts without requiring parameter updates. However, understanding how tasks are internally encoded and generalized remains a challenge. To address some of the empirical and technical gaps in the literature, we introduce an automated formulation for encoding task information in ICL prompts as a function of attention heads within the transformer architecture. This approach computes a single task vector as a weighted sum of attention heads, with the weights optimized causally via gradient descent. Our findings show that existing methods fail to generalize effectively to modalities beyond text. In response, we also design a benchmark to evaluate whether a task vector can preserve task fidelity in functional regression tasks. The proposed method successfully extracts task-specific information from in-context demonstrations and excels in both text and regression tasks, demonstrating its generalizability across modalities.
title Learning Task Representations from In-Context Learning
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2502.05390