Thumb on the Scale: Optimal Loss Weighting in Last Layer Retraining

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stromberg, Nathan, Thrampoulidis, Christos, Sankar, Lalitha
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911021229146112
author Stromberg, Nathan
Thrampoulidis, Christos
Sankar, Lalitha
author_facet Stromberg, Nathan
Thrampoulidis, Christos
Sankar, Lalitha
contents While machine learning models become more capable in discriminative tasks at scale, their ability to overcome biases introduced by training data has come under increasing scrutiny. Previous results suggest that there are two extremes of parameterization with very different behaviors: the population (underparameterized) setting where loss weighting is optimal and the separable overparameterized setting where loss weighting is ineffective at ensuring equal performance across classes. This work explores the regime of last layer retraining (LLR) in which the unseen limited (retraining) data is frequently inseparable and the model proportionately sized, falling between the two aforementioned extremes. We show, in theory and practice, that loss weighting is still effective in this regime, but that these weights \emph{must} take into account the relative overparameterization of the model.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20025
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Thumb on the Scale: Optimal Loss Weighting in Last Layer Retraining
Stromberg, Nathan
Thrampoulidis, Christos
Sankar, Lalitha
Machine Learning
While machine learning models become more capable in discriminative tasks at scale, their ability to overcome biases introduced by training data has come under increasing scrutiny. Previous results suggest that there are two extremes of parameterization with very different behaviors: the population (underparameterized) setting where loss weighting is optimal and the separable overparameterized setting where loss weighting is ineffective at ensuring equal performance across classes. This work explores the regime of last layer retraining (LLR) in which the unseen limited (retraining) data is frequently inseparable and the model proportionately sized, falling between the two aforementioned extremes. We show, in theory and practice, that loss weighting is still effective in this regime, but that these weights \emph{must} take into account the relative overparameterization of the model.
title Thumb on the Scale: Optimal Loss Weighting in Last Layer Retraining
topic Machine Learning
url https://arxiv.org/abs/2506.20025