Vector-ICL: In-context Learning with Continuous Vector Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhuang, Yufan, Singh, Chandan, Liu, Liyuan, Shang, Jingbo, Gao, Jianfeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916621150322688
author Zhuang, Yufan
Singh, Chandan
Liu, Liyuan
Shang, Jingbo
Gao, Jianfeng
author_facet Zhuang, Yufan
Singh, Chandan
Liu, Liyuan
Shang, Jingbo
Gao, Jianfeng
contents Large language models (LLMs) have shown remarkable in-context learning (ICL) capabilities on textual data. We explore whether these capabilities can be extended to continuous vectors from diverse domains, obtained from black-box pretrained encoders. By aligning input data with an LLM's embedding space through lightweight projectors, we observe that LLMs can effectively process and learn from these projected vectors, which we term Vector-ICL. In particular, we find that pretraining projectors with general language modeling objectives enables Vector-ICL, while task-specific finetuning further enhances performance. In our experiments across various tasks and modalities, including text reconstruction, numerical function regression, text classification, summarization, molecule captioning, time-series classification, graph classification, and fMRI decoding, Vector-ICL often surpasses both few-shot ICL and domain-specific model or tuning. We further conduct analyses and case studies, indicating the potential of LLMs to process vector representations beyond traditional token-based paradigms.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05629
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Vector-ICL: In-context Learning with Continuous Vector Representations
Zhuang, Yufan
Singh, Chandan
Liu, Liyuan
Shang, Jingbo
Gao, Jianfeng
Computation and Language
Artificial Intelligence
Large language models (LLMs) have shown remarkable in-context learning (ICL) capabilities on textual data. We explore whether these capabilities can be extended to continuous vectors from diverse domains, obtained from black-box pretrained encoders. By aligning input data with an LLM's embedding space through lightweight projectors, we observe that LLMs can effectively process and learn from these projected vectors, which we term Vector-ICL. In particular, we find that pretraining projectors with general language modeling objectives enables Vector-ICL, while task-specific finetuning further enhances performance. In our experiments across various tasks and modalities, including text reconstruction, numerical function regression, text classification, summarization, molecule captioning, time-series classification, graph classification, and fMRI decoding, Vector-ICL often surpasses both few-shot ICL and domain-specific model or tuning. We further conduct analyses and case studies, indicating the potential of LLMs to process vector representations beyond traditional token-based paradigms.
title Vector-ICL: In-context Learning with Continuous Vector Representations
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.05629