Unsupervised Domain Adaptation for Binary Classification with an Unobservable Source Subpopulation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866908956093317120 |
|---|---|
| author | Ying, Chao Jin, Jun Zhang, Haotian Tian, Qinglong Ma, Yanyuan Li, Sharon Zhao, Jiwei |
| author_facet | Ying, Chao Jin, Jun Zhang, Haotian Tian, Qinglong Ma, Yanyuan Li, Sharon Zhao, Jiwei |
| contents | We study an unsupervised domain adaptation problem where the source domain consists of subpopulations defined by the binary label $Y$ and a binary background (or environment) $A$. We focus on a challenging setting in which one such subpopulation in the source domain is unobservable. Naively ignoring this unobserved group can result in biased estimates and degraded predictive performance. Despite this structured missingness, we show that the prediction in the target domain can still be recovered. Specifically, we rigorously derive both background-specific and overall prediction models for the target domain. For practical implementation, we propose the distribution matching method to estimate the subpopulation proportions. We provide theoretical guarantees for the asymptotic behavior of our estimator, and establish an upper bound on the prediction error. Experiments on both synthetic and real-world datasets show that our method outperforms the naive benchmark that does not account for this unobservable source subpopulation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_20587 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Unsupervised Domain Adaptation for Binary Classification with an Unobservable Source Subpopulation Ying, Chao Jin, Jun Zhang, Haotian Tian, Qinglong Ma, Yanyuan Li, Sharon Zhao, Jiwei Machine Learning Methodology We study an unsupervised domain adaptation problem where the source domain consists of subpopulations defined by the binary label $Y$ and a binary background (or environment) $A$. We focus on a challenging setting in which one such subpopulation in the source domain is unobservable. Naively ignoring this unobserved group can result in biased estimates and degraded predictive performance. Despite this structured missingness, we show that the prediction in the target domain can still be recovered. Specifically, we rigorously derive both background-specific and overall prediction models for the target domain. For practical implementation, we propose the distribution matching method to estimate the subpopulation proportions. We provide theoretical guarantees for the asymptotic behavior of our estimator, and establish an upper bound on the prediction error. Experiments on both synthetic and real-world datasets show that our method outperforms the naive benchmark that does not account for this unobservable source subpopulation. |
| title | Unsupervised Domain Adaptation for Binary Classification with an Unobservable Source Subpopulation |
| topic | Machine Learning Methodology |
| url | https://arxiv.org/abs/2509.20587 |