Instance-Level Data-Use Auditing of Visual ML Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Zonghao, Gong, Neil Zhenqiang, Reiter, Michael K.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911156018348032
author Huang, Zonghao
Gong, Neil Zhenqiang
Reiter, Michael K.
author_facet Huang, Zonghao
Gong, Neil Zhenqiang
Reiter, Michael K.
contents The growing trend of legal disputes over the unauthorized use of data in machine learning (ML) systems highlights the urgent need for reliable data-use auditing mechanisms to ensure accountability and transparency in ML. We present the first proactive, instance-level, data-use auditing method designed to enable data owners to audit the use of their individual data instances in ML models, providing more fine-grained auditing results than previous work. To do so, our research generalizes previous work integrating black-box membership inference and sequential hypothesis testing, expanding its scope of application while preserving the quantifiable and tunable false-detection rate that is its hallmark. We evaluate our method on three types of visual ML models: image classifiers, visual encoders, and vision-language models (Contrastive Language-Image Pretraining (CLIP) and Bootstrapping Language-Image Pretraining (BLIP) models). In addition, we apply our method to evaluate the performance of two state-of-the-art approximate unlearning methods. As a noteworthy second contribution, our work reveals that neither method successfully removes the influence of the unlearned data instances from image classifiers and CLIP models, even if sacrificing model utility by $10\%$.
format Preprint
id arxiv_https___arxiv_org_abs_2503_22413
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Instance-Level Data-Use Auditing of Visual ML Models
Huang, Zonghao
Gong, Neil Zhenqiang
Reiter, Michael K.
Cryptography and Security
Machine Learning
The growing trend of legal disputes over the unauthorized use of data in machine learning (ML) systems highlights the urgent need for reliable data-use auditing mechanisms to ensure accountability and transparency in ML. We present the first proactive, instance-level, data-use auditing method designed to enable data owners to audit the use of their individual data instances in ML models, providing more fine-grained auditing results than previous work. To do so, our research generalizes previous work integrating black-box membership inference and sequential hypothesis testing, expanding its scope of application while preserving the quantifiable and tunable false-detection rate that is its hallmark. We evaluate our method on three types of visual ML models: image classifiers, visual encoders, and vision-language models (Contrastive Language-Image Pretraining (CLIP) and Bootstrapping Language-Image Pretraining (BLIP) models). In addition, we apply our method to evaluate the performance of two state-of-the-art approximate unlearning methods. As a noteworthy second contribution, our work reveals that neither method successfully removes the influence of the unlearned data instances from image classifiers and CLIP models, even if sacrificing model utility by $10\%$.
title Instance-Level Data-Use Auditing of Visual ML Models
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2503.22413