Large (Vision) Language Models are Unsupervised In-Context Learners

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gadetsky, Artyom, Atanov, Andrei, Jiang, Yulun, Gao, Zhitong, Mighan, Ghazal Hosseini, Zamir, Amir, Brbic, Maria
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912307063291904
author Gadetsky, Artyom
Atanov, Andrei
Jiang, Yulun
Gao, Zhitong
Mighan, Ghazal Hosseini
Zamir, Amir
Brbic, Maria
author_facet Gadetsky, Artyom
Atanov, Andrei
Jiang, Yulun
Gao, Zhitong
Mighan, Ghazal Hosseini
Zamir, Amir
Brbic, Maria
contents Recent advances in large language and vision-language models have enabled zero-shot inference, allowing models to solve new tasks without task-specific training. Various adaptation techniques such as prompt engineering, In-Context Learning (ICL), and supervised fine-tuning can further enhance the model's performance on a downstream task, but they require substantial manual effort to construct effective prompts or labeled examples. In this work, we introduce a joint inference framework for fully unsupervised adaptation, eliminating the need for manual prompt engineering and labeled examples. Unlike zero-shot inference, which makes independent predictions, the joint inference makes predictions simultaneously for all inputs in a given task. Since direct joint inference involves computationally expensive optimization, we develop efficient approximation techniques, leading to two unsupervised adaptation methods: unsupervised fine-tuning and unsupervised ICL. We demonstrate the effectiveness of our methods across diverse tasks and models, including language-only Llama-3.1 on natural language processing tasks, reasoning-oriented Qwen2.5-Math on grade school math problems, vision-language OpenFlamingo on vision tasks, and the API-only access GPT-4o model on massive multi-discipline tasks. Our experiments demonstrate substantial improvements over the standard zero-shot approach, including 39% absolute improvement on the challenging GSM8K math reasoning dataset. Remarkably, despite being fully unsupervised, our framework often performs on par with supervised approaches that rely on ground truth labels.
format Preprint
id arxiv_https___arxiv_org_abs_2504_02349
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large (Vision) Language Models are Unsupervised In-Context Learners
Gadetsky, Artyom
Atanov, Andrei
Jiang, Yulun
Gao, Zhitong
Mighan, Ghazal Hosseini
Zamir, Amir
Brbic, Maria
Machine Learning
Recent advances in large language and vision-language models have enabled zero-shot inference, allowing models to solve new tasks without task-specific training. Various adaptation techniques such as prompt engineering, In-Context Learning (ICL), and supervised fine-tuning can further enhance the model's performance on a downstream task, but they require substantial manual effort to construct effective prompts or labeled examples. In this work, we introduce a joint inference framework for fully unsupervised adaptation, eliminating the need for manual prompt engineering and labeled examples. Unlike zero-shot inference, which makes independent predictions, the joint inference makes predictions simultaneously for all inputs in a given task. Since direct joint inference involves computationally expensive optimization, we develop efficient approximation techniques, leading to two unsupervised adaptation methods: unsupervised fine-tuning and unsupervised ICL. We demonstrate the effectiveness of our methods across diverse tasks and models, including language-only Llama-3.1 on natural language processing tasks, reasoning-oriented Qwen2.5-Math on grade school math problems, vision-language OpenFlamingo on vision tasks, and the API-only access GPT-4o model on massive multi-discipline tasks. Our experiments demonstrate substantial improvements over the standard zero-shot approach, including 39% absolute improvement on the challenging GSM8K math reasoning dataset. Remarkably, despite being fully unsupervised, our framework often performs on par with supervised approaches that rely on ground truth labels.
title Large (Vision) Language Models are Unsupervised In-Context Learners
topic Machine Learning
url https://arxiv.org/abs/2504.02349