Aggregating Data for Optimal and Private Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Agarwal, Sushant, Makhija, Yukti, Saket, Rishi, Raghuveer, Aravindan
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929608442511360
author Agarwal, Sushant
Makhija, Yukti
Saket, Rishi
Raghuveer, Aravindan
author_facet Agarwal, Sushant
Makhija, Yukti
Saket, Rishi
Raghuveer, Aravindan
contents Multiple Instance Regression (MIR) and Learning from Label Proportions (LLP) are learning frameworks arising in many applications, where the training data is partitioned into disjoint sets or bags, and only an aggregate label i.e., bag-label for each bag is available to the learner. In the case of MIR, the bag-label is the label of an undisclosed instance from the bag, while in LLP, the bag-label is the mean of the bag's labels. In this paper, we study for various loss functions in MIR and LLP, what is the optimal way to partition the dataset into bags such that the utility for downstream tasks like linear regression is maximized. We theoretically provide utility guarantees, and show that in each case, the optimal bagging strategy (approximately) reduces to finding an optimal clustering of the feature vectors or the labels with respect to natural objectives such as $k$-means. We also show that our bagging mechanisms can be made label-differentially private, incurring an additional utility error. We then generalize our results to the setting of Generalized Linear Models (GLMs). Finally, we experimentally validate our theoretical results.
format Preprint
id arxiv_https___arxiv_org_abs_2411_19045
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Aggregating Data for Optimal and Private Learning
Agarwal, Sushant
Makhija, Yukti
Saket, Rishi
Raghuveer, Aravindan
Machine Learning
Multiple Instance Regression (MIR) and Learning from Label Proportions (LLP) are learning frameworks arising in many applications, where the training data is partitioned into disjoint sets or bags, and only an aggregate label i.e., bag-label for each bag is available to the learner. In the case of MIR, the bag-label is the label of an undisclosed instance from the bag, while in LLP, the bag-label is the mean of the bag's labels. In this paper, we study for various loss functions in MIR and LLP, what is the optimal way to partition the dataset into bags such that the utility for downstream tasks like linear regression is maximized. We theoretically provide utility guarantees, and show that in each case, the optimal bagging strategy (approximately) reduces to finding an optimal clustering of the feature vectors or the labels with respect to natural objectives such as $k$-means. We also show that our bagging mechanisms can be made label-differentially private, incurring an additional utility error. We then generalize our results to the setting of Generalized Linear Models (GLMs). Finally, we experimentally validate our theoretical results.
title Aggregating Data for Optimal and Private Learning
topic Machine Learning
url https://arxiv.org/abs/2411.19045