Metric-DST: Mitigating Selection Bias Through Diversity-Guided Semi-Supervised Metric Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929608268447744 |
|---|---|
| author | Tepeli, Yasin I. de Wolf, Mathijs Gonçalves, Joana P. |
| author_facet | Tepeli, Yasin I. de Wolf, Mathijs Gonçalves, Joana P. |
| contents | Selection bias poses a critical challenge for fairness in machine learning, as models trained on data that is less representative of the population might exhibit undesirable behavior for underrepresented profiles. Semi-supervised learning strategies like self-training can mitigate selection bias by incorporating unlabeled data into model training to gain further insight into the distribution of the population. However, conventional self-training seeks to include high-confidence data samples, which may reinforce existing model bias and compromise effectiveness. We propose Metric-DST, a diversity-guided self-training strategy that leverages metric learning and its implicit embedding space to counter confidence-based bias through the inclusion of more diverse samples. Metric-DST learned more robust models in the presence of selection bias for generated and real-world datasets with induced bias, as well as a molecular biology prediction task with intrinsic bias. The Metric-DST learning strategy offers a flexible and widely applicable solution to mitigate selection bias and enhance fairness of machine learning models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_18442 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Metric-DST: Mitigating Selection Bias Through Diversity-Guided Semi-Supervised Metric Learning Tepeli, Yasin I. de Wolf, Mathijs Gonçalves, Joana P. Machine Learning Artificial Intelligence Selection bias poses a critical challenge for fairness in machine learning, as models trained on data that is less representative of the population might exhibit undesirable behavior for underrepresented profiles. Semi-supervised learning strategies like self-training can mitigate selection bias by incorporating unlabeled data into model training to gain further insight into the distribution of the population. However, conventional self-training seeks to include high-confidence data samples, which may reinforce existing model bias and compromise effectiveness. We propose Metric-DST, a diversity-guided self-training strategy that leverages metric learning and its implicit embedding space to counter confidence-based bias through the inclusion of more diverse samples. Metric-DST learned more robust models in the presence of selection bias for generated and real-world datasets with induced bias, as well as a molecular biology prediction task with intrinsic bias. The Metric-DST learning strategy offers a flexible and widely applicable solution to mitigate selection bias and enhance fairness of machine learning models. |
| title | Metric-DST: Mitigating Selection Bias Through Diversity-Guided Semi-Supervised Metric Learning |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2411.18442 |