Déjà Vu Memorization in Vision-Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jayaraman, Bargav, Guo, Chuan, Chaudhuri, Kamalika
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917819416838144
author Jayaraman, Bargav
Guo, Chuan
Chaudhuri, Kamalika
author_facet Jayaraman, Bargav
Guo, Chuan
Chaudhuri, Kamalika
contents Vision-Language Models (VLMs) have emerged as the state-of-the-art representation learning solution, with myriads of downstream applications such as image classification, retrieval and generation. A natural question is whether these models memorize their training data, which also has implications for generalization. We propose a new method for measuring memorization in VLMs, which we call déjà vu memorization. For VLMs trained on image-caption pairs, we show that the model indeed retains information about individual objects in the training images beyond what can be inferred from correlations or the image caption. We evaluate déjà vu memorization at both sample and population level, and show that it is significant for OpenCLIP trained on as many as 50M image-caption pairs. Finally, we show that text randomization considerably mitigates memorization while only moderately impacting the model's downstream task performance.
format Preprint
id arxiv_https___arxiv_org_abs_2402_02103
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Déjà Vu Memorization in Vision-Language Models
Jayaraman, Bargav
Guo, Chuan
Chaudhuri, Kamalika
Computer Vision and Pattern Recognition
Machine Learning
Vision-Language Models (VLMs) have emerged as the state-of-the-art representation learning solution, with myriads of downstream applications such as image classification, retrieval and generation. A natural question is whether these models memorize their training data, which also has implications for generalization. We propose a new method for measuring memorization in VLMs, which we call déjà vu memorization. For VLMs trained on image-caption pairs, we show that the model indeed retains information about individual objects in the training images beyond what can be inferred from correlations or the image caption. We evaluate déjà vu memorization at both sample and population level, and show that it is significant for OpenCLIP trained on as many as 50M image-caption pairs. Finally, we show that text randomization considerably mitigates memorization while only moderately impacting the model's downstream task performance.
title Déjà Vu Memorization in Vision-Language Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2402.02103