Fine-Tuned In-Context Learners for Efficient Adaptation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Bornschein, Jorg, Lyle, Clare, Li, Yazhe, Rannen-Triki, Amal, He, Xu Owen, Pascanu, Razvan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908729886113792
author Bornschein, Jorg
Lyle, Clare
Li, Yazhe
Rannen-Triki, Amal
He, Xu Owen
Pascanu, Razvan
author_facet Bornschein, Jorg
Lyle, Clare
Li, Yazhe
Rannen-Triki, Amal
He, Xu Owen
Pascanu, Razvan
contents When adapting large language models (LLMs) to a specific downstream task, two primary approaches are commonly employed: (1) prompt engineering, often with in-context few-shot learning, leveraging the model's inherent generalization abilities, and (2) fine-tuning on task-specific data, directly optimizing the model's parameters. While prompt-based methods excel in few-shot scenarios, their effectiveness often plateaus as more data becomes available. Conversely, fine-tuning scales well with data but may underperform when training examples are scarce. We investigate a unified approach that bridges these two paradigms by incorporating in-context learning directly into the fine-tuning process. Specifically, we fine-tune the model on task-specific data augmented with in-context examples, mimicking the structure of k-shot prompts. This approach, while requiring per-task fine-tuning, combines the sample efficiency of in-context learning with the performance gains of fine-tuning, leading to a method that consistently matches and often significantly exceeds both these baselines. To perform hyperparameter selection in the low-data regime, we propose to use prequential evaluation, which eliminates the need for expensive cross-validation and leverages all available data for training while simultaneously providing a robust validation signal. We conduct an extensive empirical study to determine which adaptation paradigm - fine-tuning, in-context learning, or our proposed unified approach offers the best predictive performance on a concrete data downstream-tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2512_19879
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-Tuned In-Context Learners for Efficient Adaptation
Bornschein, Jorg
Lyle, Clare
Li, Yazhe
Rannen-Triki, Amal
He, Xu Owen
Pascanu, Razvan
Machine Learning
Artificial Intelligence
When adapting large language models (LLMs) to a specific downstream task, two primary approaches are commonly employed: (1) prompt engineering, often with in-context few-shot learning, leveraging the model's inherent generalization abilities, and (2) fine-tuning on task-specific data, directly optimizing the model's parameters. While prompt-based methods excel in few-shot scenarios, their effectiveness often plateaus as more data becomes available. Conversely, fine-tuning scales well with data but may underperform when training examples are scarce. We investigate a unified approach that bridges these two paradigms by incorporating in-context learning directly into the fine-tuning process. Specifically, we fine-tune the model on task-specific data augmented with in-context examples, mimicking the structure of k-shot prompts. This approach, while requiring per-task fine-tuning, combines the sample efficiency of in-context learning with the performance gains of fine-tuning, leading to a method that consistently matches and often significantly exceeds both these baselines. To perform hyperparameter selection in the low-data regime, we propose to use prequential evaluation, which eliminates the need for expensive cross-validation and leverages all available data for training while simultaneously providing a robust validation signal. We conduct an extensive empirical study to determine which adaptation paradigm - fine-tuning, in-context learning, or our proposed unified approach offers the best predictive performance on a concrete data downstream-tasks.
title Fine-Tuned In-Context Learners for Efficient Adaptation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2512.19879