Pseudo-Labeling for Unsupervised Domain Adaptation with Kernel GLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Weill, Nathan, Wang, Kaizheng
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908906597384192
author Weill, Nathan
Wang, Kaizheng
author_facet Weill, Nathan
Wang, Kaizheng
contents We propose a principled framework for unsupervised domain adaptation under covariate shift in kernel Generalized Linear Models (GLMs), encompassing kernelized linear, logistic, and Poisson regression with ridge regularization. Our goal is to minimize prediction error in the target domain by leveraging labeled source data and unlabeled target data, despite differences in covariate distributions. We partition the labeled source data into two batches: one for training a family of candidate models, and the other for building an imputation model. This imputation model generates pseudo-labels for the target data, enabling robust model selection. We establish non-asymptotic excess-risk bounds that characterize adaptation performance through an "effective labeled sample size", explicitly accounting for the unknown covariate shift. Experiments on synthetic and real datasets demonstrate consistent performance gains over source-only baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19422
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Pseudo-Labeling for Unsupervised Domain Adaptation with Kernel GLMs
Weill, Nathan
Wang, Kaizheng
Machine Learning
Statistics Theory
We propose a principled framework for unsupervised domain adaptation under covariate shift in kernel Generalized Linear Models (GLMs), encompassing kernelized linear, logistic, and Poisson regression with ridge regularization. Our goal is to minimize prediction error in the target domain by leveraging labeled source data and unlabeled target data, despite differences in covariate distributions. We partition the labeled source data into two batches: one for training a family of candidate models, and the other for building an imputation model. This imputation model generates pseudo-labels for the target data, enabling robust model selection. We establish non-asymptotic excess-risk bounds that characterize adaptation performance through an "effective labeled sample size", explicitly accounting for the unknown covariate shift. Experiments on synthetic and real datasets demonstrate consistent performance gains over source-only baselines.
title Pseudo-Labeling for Unsupervised Domain Adaptation with Kernel GLMs
topic Machine Learning
Statistics Theory
url https://arxiv.org/abs/2603.19422