Mitigating covariate shift in non-colocated data with learned parameter priors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khan, Behraj, Mirza, Behroz, Durrani, Nouman, Syed, Tahir
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913575533019136
author Khan, Behraj
Mirza, Behroz
Durrani, Nouman
Syed, Tahir
author_facet Khan, Behraj
Mirza, Behroz
Durrani, Nouman
Syed, Tahir
contents When training data are distributed across{ time or space,} covariate shift across fragments of training data biases cross-validation, compromising model selection and assessment. We present \textit{Fragmentation-Induced covariate-shift Remediation} ($FIcsR$), which minimizes an $f$-divergence between a fragment's covariate distribution and that of the standard cross-validation baseline. We s{how} an equivalence with popular importance-weighting methods. {The method}'s numerical solution poses a computational challenge owing to the overparametrized nature of a neural network, and we derive a Fisher Information approximation. When accumulated over fragments, this provides a global estimate of the amount of shift remediation thus far needed, and we incorporate that as a prior via the minimization objective. In the paper, we run extensive classification experiments on multiple data classes, over $40$ datasets, and with data batched over multiple sequence lengths. We extend the study to the $k$-fold cross-validation setting through a similar set of experiments. An ablation study exposes the method to varying amounts of shift and demonstrates slower degradation with $FIcsR$ in place. The results are promising under all these conditions; with improved accuracy against batch and fold state-of-the-art by more than $5\%$ and $10\%$, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2411_06499
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mitigating covariate shift in non-colocated data with learned parameter priors
Khan, Behraj
Mirza, Behroz
Durrani, Nouman
Syed, Tahir
Machine Learning
Computer Vision and Pattern Recognition
When training data are distributed across{ time or space,} covariate shift across fragments of training data biases cross-validation, compromising model selection and assessment. We present \textit{Fragmentation-Induced covariate-shift Remediation} ($FIcsR$), which minimizes an $f$-divergence between a fragment's covariate distribution and that of the standard cross-validation baseline. We s{how} an equivalence with popular importance-weighting methods. {The method}'s numerical solution poses a computational challenge owing to the overparametrized nature of a neural network, and we derive a Fisher Information approximation. When accumulated over fragments, this provides a global estimate of the amount of shift remediation thus far needed, and we incorporate that as a prior via the minimization objective. In the paper, we run extensive classification experiments on multiple data classes, over $40$ datasets, and with data batched over multiple sequence lengths. We extend the study to the $k$-fold cross-validation setting through a similar set of experiments. An ablation study exposes the method to varying amounts of shift and demonstrates slower degradation with $FIcsR$ in place. The results are promising under all these conditions; with improved accuracy against batch and fold state-of-the-art by more than $5\%$ and $10\%$, respectively.
title Mitigating covariate shift in non-colocated data with learned parameter priors
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.06499