The Magic Correlations: Understanding Knowledge Transfer from Pretraining to Supervised Fine-Tuning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Fan, Simin, Paparas, Dimitris, Noy, Natasha, Xiong, Binbin, Sachdeva, Noveen, Isik, Berivan
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908828723838976
author Fan, Simin
Paparas, Dimitris
Noy, Natasha
Xiong, Binbin
Sachdeva, Noveen
Isik, Berivan
author_facet Fan, Simin
Paparas, Dimitris
Noy, Natasha
Xiong, Binbin
Sachdeva, Noveen
Isik, Berivan
contents Understanding how language model capabilities transfer from pretraining to supervised fine-tuning (SFT) is fundamental to efficient model development and data curation. In this work, we investigate four core questions: RQ1. To what extent do accuracy and confidence rankings established during pretraining persist after SFT? RQ2. Which benchmarks serve as robust cross-stage predictors and which are unreliable? RQ3. How do transfer dynamics shift with model scale? RQ4. How well does model confidence align with accuracy, as a measure of calibration quality? Does this alignment pattern transfer across training stages? We address these questions through a suite of correlation protocols applied to accuracy and confidence metrics across diverse data mixtures and model scales. Our experiments reveal that transfer reliability varies dramatically across capability categories, benchmarks, and scales -- with accuracy and confidence exhibiting distinct, sometimes opposing, scaling dynamics. These findings shed light on the complex interplay between pretraining decisions and downstream outcomes, providing actionable guidance for benchmark selection, data curation, and efficient model development.
format Preprint
id arxiv_https___arxiv_org_abs_2602_11217
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Magic Correlations: Understanding Knowledge Transfer from Pretraining to Supervised Fine-Tuning
Fan, Simin
Paparas, Dimitris
Noy, Natasha
Xiong, Binbin
Sachdeva, Noveen
Isik, Berivan
Machine Learning
Understanding how language model capabilities transfer from pretraining to supervised fine-tuning (SFT) is fundamental to efficient model development and data curation. In this work, we investigate four core questions: RQ1. To what extent do accuracy and confidence rankings established during pretraining persist after SFT? RQ2. Which benchmarks serve as robust cross-stage predictors and which are unreliable? RQ3. How do transfer dynamics shift with model scale? RQ4. How well does model confidence align with accuracy, as a measure of calibration quality? Does this alignment pattern transfer across training stages? We address these questions through a suite of correlation protocols applied to accuracy and confidence metrics across diverse data mixtures and model scales. Our experiments reveal that transfer reliability varies dramatically across capability categories, benchmarks, and scales -- with accuracy and confidence exhibiting distinct, sometimes opposing, scaling dynamics. These findings shed light on the complex interplay between pretraining decisions and downstream outcomes, providing actionable guidance for benchmark selection, data curation, and efficient model development.
title The Magic Correlations: Understanding Knowledge Transfer from Pretraining to Supervised Fine-Tuning
topic Machine Learning
url https://arxiv.org/abs/2602.11217