Revisiting Data Auditing in Large Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Hongyu, Liang, Sichu, Wang, Wenwen, Li, Boheng, Yuan, Tongxin, Li, Fangqi, Wang, ShiLin, Zhang, Zhuosheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915258753351680
author Zhu, Hongyu
Liang, Sichu
Wang, Wenwen
Li, Boheng
Yuan, Tongxin
Li, Fangqi
Wang, ShiLin
Zhang, Zhuosheng
author_facet Zhu, Hongyu
Liang, Sichu
Wang, Wenwen
Li, Boheng
Yuan, Tongxin
Li, Fangqi
Wang, ShiLin
Zhang, Zhuosheng
contents With the surge of large language models (LLMs), Large Vision-Language Models (VLMs)--which integrate vision encoders with LLMs for accurate visual grounding--have shown great potential in tasks like generalist agents and robotic control. However, VLMs are typically trained on massive web-scraped images, raising concerns over copyright infringement and privacy violations, and making data auditing increasingly urgent. Membership inference (MI), which determines whether a sample was used in training, has emerged as a key auditing technique, with promising results on open-source VLMs like LLaVA (AUC > 80%). In this work, we revisit these advances and uncover a critical issue: current MI benchmarks suffer from distribution shifts between member and non-member images, introducing shortcut cues that inflate MI performance. We further analyze the nature of these shifts and propose a principled metric based on optimal transport to quantify the distribution discrepancy. To evaluate MI in realistic settings, we construct new benchmarks with i.i.d. member and non-member images. Existing MI methods fail under these unbiased conditions, performing only marginally better than chance. Further, we explore the theoretical upper bound of MI by probing the Bayes Optimality within the VLM's embedding space and find the irreducible error rate remains high. Despite this pessimistic outlook, we analyze why MI for VLMs is particularly challenging and identify three practical scenarios--fine-tuning, access to ground-truth texts, and set-based inference--where auditing becomes feasible. Our study presents a systematic view of the limits and opportunities of MI for VLMs, providing guidance for future efforts in trustworthy data auditing.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18349
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revisiting Data Auditing in Large Vision-Language Models
Zhu, Hongyu
Liang, Sichu
Wang, Wenwen
Li, Boheng
Yuan, Tongxin
Li, Fangqi
Wang, ShiLin
Zhang, Zhuosheng
Computer Vision and Pattern Recognition
Cryptography and Security
With the surge of large language models (LLMs), Large Vision-Language Models (VLMs)--which integrate vision encoders with LLMs for accurate visual grounding--have shown great potential in tasks like generalist agents and robotic control. However, VLMs are typically trained on massive web-scraped images, raising concerns over copyright infringement and privacy violations, and making data auditing increasingly urgent. Membership inference (MI), which determines whether a sample was used in training, has emerged as a key auditing technique, with promising results on open-source VLMs like LLaVA (AUC > 80%). In this work, we revisit these advances and uncover a critical issue: current MI benchmarks suffer from distribution shifts between member and non-member images, introducing shortcut cues that inflate MI performance. We further analyze the nature of these shifts and propose a principled metric based on optimal transport to quantify the distribution discrepancy. To evaluate MI in realistic settings, we construct new benchmarks with i.i.d. member and non-member images. Existing MI methods fail under these unbiased conditions, performing only marginally better than chance. Further, we explore the theoretical upper bound of MI by probing the Bayes Optimality within the VLM's embedding space and find the irreducible error rate remains high. Despite this pessimistic outlook, we analyze why MI for VLMs is particularly challenging and identify three practical scenarios--fine-tuning, access to ground-truth texts, and set-based inference--where auditing becomes feasible. Our study presents a systematic view of the limits and opportunities of MI for VLMs, providing guidance for future efforts in trustworthy data auditing.
title Revisiting Data Auditing in Large Vision-Language Models
topic Computer Vision and Pattern Recognition
Cryptography and Security
url https://arxiv.org/abs/2504.18349