Beyond Pooling: Matching for Robust Generalization under Data Heterogeneity

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Roy, Ayush, Chakraborty, Rudrasis, Varshney, Lav, Lokhande, Vishnu Suresh
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911428679565312
author Roy, Ayush
Chakraborty, Rudrasis
Varshney, Lav
Lokhande, Vishnu Suresh
author_facet Roy, Ayush
Chakraborty, Rudrasis
Varshney, Lav
Lokhande, Vishnu Suresh
contents Pooling heterogeneous datasets across domains is a common strategy in representation learning, but naive pooling can amplify distributional asymmetries and yield biased estimators, especially in settings where zero-shot generalization is required. We propose a matching framework that selects samples relative to an adaptive centroid and iteratively refines the representation distribution. The double robustness and the propensity score matching for the inclusion of data domains make matching more robust than naive pooling and uniform subsampling by filtering out the confounding domains (the main cause of heterogeneity). Theoretical and empirical analyses show that, unlike naive pooling or uniform subsampling, matching achieves better results under asymmetric meta-distributions, which are also extended to non-Gaussian and multimodal real-world settings. Most importantly, we show that these improvements translate to zero-shot medical anomaly detection, one of the extreme forms of data heterogeneity and asymmetry. The code is available on https://github.com/AyushRoy2001/Beyond-Pooling.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07154
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Pooling: Matching for Robust Generalization under Data Heterogeneity
Roy, Ayush
Chakraborty, Rudrasis
Varshney, Lav
Lokhande, Vishnu Suresh
Machine Learning
Artificial Intelligence
Pooling heterogeneous datasets across domains is a common strategy in representation learning, but naive pooling can amplify distributional asymmetries and yield biased estimators, especially in settings where zero-shot generalization is required. We propose a matching framework that selects samples relative to an adaptive centroid and iteratively refines the representation distribution. The double robustness and the propensity score matching for the inclusion of data domains make matching more robust than naive pooling and uniform subsampling by filtering out the confounding domains (the main cause of heterogeneity). Theoretical and empirical analyses show that, unlike naive pooling or uniform subsampling, matching achieves better results under asymmetric meta-distributions, which are also extended to non-Gaussian and multimodal real-world settings. Most importantly, we show that these improvements translate to zero-shot medical anomaly detection, one of the extreme forms of data heterogeneity and asymmetry. The code is available on https://github.com/AyushRoy2001/Beyond-Pooling.
title Beyond Pooling: Matching for Robust Generalization under Data Heterogeneity
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.07154