Observational Auditing of Label Privacy

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kalemaj, Iden, Melis, Luca, Boucher, Maxime, Mironov, Ilya, Mahloujifar, Saeed
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909992960917504
author Kalemaj, Iden
Melis, Luca
Boucher, Maxime
Mironov, Ilya
Mahloujifar, Saeed
author_facet Kalemaj, Iden
Melis, Luca
Boucher, Maxime
Mironov, Ilya
Mahloujifar, Saeed
contents Differential privacy (DP) auditing is essential for evaluating privacy guarantees in machine learning systems. Existing auditing methods, however, pose a significant challenge for large-scale systems since they require modifying the training dataset -- for instance, by injecting out-of-distribution canaries or removing samples from training. Such interventions on the training data pipeline are resource-intensive and involve considerable engineering overhead. We introduce a novel observational auditing framework that leverages the inherent randomness of data distributions, enabling privacy evaluation without altering the original dataset. Our approach extends privacy auditing beyond traditional membership inference to protected attributes, with labels as a special case, addressing a key gap in existing techniques. We provide theoretical foundations for our method and perform experiments on Criteo and CIFAR-10 datasets that demonstrate its effectiveness in auditing label privacy guarantees. This work opens new avenues for practical privacy auditing in large-scale production environments.
format Preprint
id arxiv_https___arxiv_org_abs_2511_14084
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Observational Auditing of Label Privacy
Kalemaj, Iden
Melis, Luca
Boucher, Maxime
Mironov, Ilya
Mahloujifar, Saeed
Machine Learning
Cryptography and Security
Differential privacy (DP) auditing is essential for evaluating privacy guarantees in machine learning systems. Existing auditing methods, however, pose a significant challenge for large-scale systems since they require modifying the training dataset -- for instance, by injecting out-of-distribution canaries or removing samples from training. Such interventions on the training data pipeline are resource-intensive and involve considerable engineering overhead. We introduce a novel observational auditing framework that leverages the inherent randomness of data distributions, enabling privacy evaluation without altering the original dataset. Our approach extends privacy auditing beyond traditional membership inference to protected attributes, with labels as a special case, addressing a key gap in existing techniques. We provide theoretical foundations for our method and perform experiments on Criteo and CIFAR-10 datasets that demonstrate its effectiveness in auditing label privacy guarantees. This work opens new avenues for practical privacy auditing in large-scale production environments.
title Observational Auditing of Label Privacy
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2511.14084