IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mishra, Aayush, Khashabi, Daniel, Liu, Anqi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910058123624448
author Mishra, Aayush
Khashabi, Daniel
Liu, Anqi
author_facet Mishra, Aayush
Khashabi, Daniel
Liu, Anqi
contents Supervised Fine-Tuning (SFT) is used to specialize model behavior by training weights to produce intended target responses for queries. In contrast, In-Context Learning (ICL) adapts models during inference with instructions or demonstrations in the prompt. ICL can offer better generalizability and more calibrated responses compared to SFT in data scarce settings, at the cost of more inference compute. In this work, we ask the question: Can ICL's internal computations be used to improve the qualities of SFT? We first show that ICL and SFT produce distinct activation patterns, indicating that the two methods achieve adaptation through different functional mechanisms. Motivated by this observation and to use ICL's rich functionality, we introduce ICL Activation Alignment (IA2), a self-distillation technique which aims to replicate ICL's activation patterns in SFT models and incentivizes ICL-like internal reasoning. Performing IA2 as a priming step before SFT significantly improves the accuracy and calibration of model outputs, as shown by our extensive empirical results on 12 popular benchmarks and two model families. This finding is not only practically useful, but also offers a conceptual window into the inner mechanics of model adaptation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22621
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
Mishra, Aayush
Khashabi, Daniel
Liu, Anqi
Machine Learning
Artificial Intelligence
Computation and Language
Supervised Fine-Tuning (SFT) is used to specialize model behavior by training weights to produce intended target responses for queries. In contrast, In-Context Learning (ICL) adapts models during inference with instructions or demonstrations in the prompt. ICL can offer better generalizability and more calibrated responses compared to SFT in data scarce settings, at the cost of more inference compute. In this work, we ask the question: Can ICL's internal computations be used to improve the qualities of SFT? We first show that ICL and SFT produce distinct activation patterns, indicating that the two methods achieve adaptation through different functional mechanisms. Motivated by this observation and to use ICL's rich functionality, we introduce ICL Activation Alignment (IA2), a self-distillation technique which aims to replicate ICL's activation patterns in SFT models and incentivizes ICL-like internal reasoning. Performing IA2 as a priming step before SFT significantly improves the accuracy and calibration of model outputs, as shown by our extensive empirical results on 12 popular benchmarks and two model families. This finding is not only practically useful, but also offers a conceptual window into the inner mechanics of model adaptation.
title IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.22621