Foundation Model's Embedded Representations May Detect Distribution Shift

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vargas, Max, Tsou, Adam, Engel, Andrew, Chiang, Tony
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910315536449536
author Vargas, Max
Tsou, Adam
Engel, Andrew
Chiang, Tony
author_facet Vargas, Max
Tsou, Adam
Engel, Andrew
Chiang, Tony
contents Sampling biases can cause distribution shifts between train and test datasets for supervised learning tasks, obscuring our ability to understand the generalization capacity of a model. This is especially important considering the wide adoption of pre-trained foundational neural networks -- whose behavior remains poorly understood -- for transfer learning (TL) tasks. We present a case study for TL on the Sentiment140 dataset and show that many pre-trained foundation models encode different representations of Sentiment140's manually curated test set $M$ from the automatically labeled training set $P$, confirming that a distribution shift has occurred. We argue training on $P$ and measuring performance on $M$ is a biased measure of generalization. Experiments on pre-trained GPT-2 show that the features learnable from $P$ do not improve (and in fact hamper) performance on $M$. Linear probes on pre-trained GPT-2's representations are robust and may even outperform overall fine-tuning, implying a fundamental importance for discerning distribution shift in train/test splits for model interpretation.
format Preprint
id arxiv_https___arxiv_org_abs_2310_13836
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Foundation Model's Embedded Representations May Detect Distribution Shift
Vargas, Max
Tsou, Adam
Engel, Andrew
Chiang, Tony
Machine Learning
Computation and Language
Sampling biases can cause distribution shifts between train and test datasets for supervised learning tasks, obscuring our ability to understand the generalization capacity of a model. This is especially important considering the wide adoption of pre-trained foundational neural networks -- whose behavior remains poorly understood -- for transfer learning (TL) tasks. We present a case study for TL on the Sentiment140 dataset and show that many pre-trained foundation models encode different representations of Sentiment140's manually curated test set $M$ from the automatically labeled training set $P$, confirming that a distribution shift has occurred. We argue training on $P$ and measuring performance on $M$ is a biased measure of generalization. Experiments on pre-trained GPT-2 show that the features learnable from $P$ do not improve (and in fact hamper) performance on $M$. Linear probes on pre-trained GPT-2's representations are robust and may even outperform overall fine-tuning, implying a fundamental importance for discerning distribution shift in train/test splits for model interpretation.
title Foundation Model's Embedded Representations May Detect Distribution Shift
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2310.13836