Captured by Captions: On Memorization and its Mitigation in CLIP Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Wenhao, Dziedzic, Adam, Kim, Grace C., Backes, Michael, Boenisch, Franziska |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Localizing Memorization in SSL Vision Encoders
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
Finding DoRI: Discovery of Retained Images in Diffusion Models
by: Kowalczuk, Antoni, et al.
Published: (2025)
by: Kowalczuk, Antoni, et al.
Published: (2025)
Demystifying Foreground-Background Memorization in Diffusion Models
by: Di, Jimmy Z., et al.
Published: (2025)
by: Di, Jimmy Z., et al.
Published: (2025)
Privacy Attacks on Image AutoRegressive Models
by: Kowalczuk, Antoni, et al.
Published: (2025)
by: Kowalczuk, Antoni, et al.
Published: (2025)
Radioactive Watermarks in Diffusion and Autoregressive Image Generative Models
by: Meintz, Michel, et al.
Published: (2025)
by: Meintz, Michel, et al.
Published: (2025)
BitMark: Watermarking Bitwise Autoregressive Image Generative Models
by: Kerner, Louis, et al.
Published: (2025)
by: Kerner, Louis, et al.
Published: (2025)
Finding NeMo: Localizing Neurons Responsible For Memorization in Diffusion Models
by: Hintersdorf, Dominik, et al.
Published: (2024)
by: Hintersdorf, Dominik, et al.
Published: (2024)
Localizing and Mitigating Memorization in Image Autoregressive Models
by: Kasliwal, Aditya, et al.
Published: (2025)
by: Kasliwal, Aditya, et al.
Published: (2025)
Memorization In Stable Diffusion Is Unexpectedly Driven by CLIP Embeddings
by: Kim, Bumjun, et al.
Published: (2026)
by: Kim, Bumjun, et al.
Published: (2026)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
by: Lai, Zhengfeng, et al.
Published: (2023)
by: Lai, Zhengfeng, et al.
Published: (2023)
Conditioned Activation Transport for T2I Safety Steering
by: Chrabąszcz, Maciej, et al.
Published: (2026)
by: Chrabąszcz, Maciej, et al.
Published: (2026)
Implementing Adaptations for Vision AutoRegressive Model
by: Shaikh, Kaif, et al.
Published: (2025)
by: Shaikh, Kaif, et al.
Published: (2025)
Detecting and Mitigating Memorization in Diffusion Models through Anisotropy of the Log-Probability
by: Asthana, Rohan, et al.
Published: (2026)
by: Asthana, Rohan, et al.
Published: (2026)
Memorization in Self-Supervised Learning Improves Downstream Generalization
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
On Memorization in Diffusion Models
by: Gu, Xiangming, et al.
Published: (2023)
by: Gu, Xiangming, et al.
Published: (2023)
Unlocking Post-hoc Dataset Inference with Synthetic Data
by: Zhao, Bihe, et al.
Published: (2025)
by: Zhao, Bihe, et al.
Published: (2025)
CLIP Can Understand Depth
by: Kim, Sohee, et al.
Published: (2024)
by: Kim, Sohee, et al.
Published: (2024)
SafeR-CLIP: Mitigating NSFW Content in Vision-Language Models While Preserving Pre-Trained Knowledge
by: Yousaf, Adeel, et al.
Published: (2025)
by: Yousaf, Adeel, et al.
Published: (2025)
Zero-Shot, But at What Cost? Unveiling the Hidden Overhead of MILS's LLM-CLIP Framework for Image Captioning
by: Benhammou, Yassir, et al.
Published: (2025)
by: Benhammou, Yassir, et al.
Published: (2025)
Benchmarking Robust Self-Supervised Learning Across Diverse Downstream Tasks
by: Kowalczuk, Antoni, et al.
Published: (2024)
by: Kowalczuk, Antoni, et al.
Published: (2024)
DiffCLIP: Differential Attention Meets CLIP
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
Memorization in Graph Neural Networks
by: Jamadandi, Adarsh, et al.
Published: (2025)
by: Jamadandi, Adarsh, et al.
Published: (2025)
CountCLIP -- [Re] Teaching CLIP to Count to Ten
by: Mestha, Harshvardhan, et al.
Published: (2024)
by: Mestha, Harshvardhan, et al.
Published: (2024)
Generalizable Geometric Image Caption Synthesis
by: Xin, Yue, et al.
Published: (2025)
by: Xin, Yue, et al.
Published: (2025)
Steering Away from Memorization: Reachability-Constrained Reinforcement Learning for Text-to-Image Diffusion
by: Karnik, Sathwik, et al.
Published: (2026)
by: Karnik, Sathwik, et al.
Published: (2026)
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
by: Fiastre, Gabriel, et al.
Published: (2025)
by: Fiastre, Gabriel, et al.
Published: (2025)
Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions
by: Hsieh, Yu-Guan, et al.
Published: (2024)
by: Hsieh, Yu-Guan, et al.
Published: (2024)
Image Captions are Natural Prompts for Text-to-Image Models
by: Lei, Shiye, et al.
Published: (2023)
by: Lei, Shiye, et al.
Published: (2023)
Memorized Images in Diffusion Models share a Subspace that can be Located and Deleted
by: Chavhan, Ruchika, et al.
Published: (2024)
by: Chavhan, Ruchika, et al.
Published: (2024)
Beautiful Images, Toxic Words: Understanding and Addressing Offensive Text in Generated Images
by: Kumar, Aditya, et al.
Published: (2025)
by: Kumar, Aditya, et al.
Published: (2025)
Impact of Layer Norm on Memorization and Generalization in Transformers
by: Singhal, Rishi, et al.
Published: (2025)
by: Singhal, Rishi, et al.
Published: (2025)
CompCap: Improving Multimodal Large Language Models with Composite Captions
by: Chen, Xiaohui, et al.
Published: (2024)
by: Chen, Xiaohui, et al.
Published: (2024)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models
by: Kao, Kuei-Chun, et al.
Published: (2025)
by: Kao, Kuei-Chun, et al.
Published: (2025)
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
by: Chu, Tianzhe, et al.
Published: (2025)
by: Chu, Tianzhe, et al.
Published: (2025)
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
by: Kim, Si-Woo, et al.
Published: (2025)
by: Kim, Si-Woo, et al.
Published: (2025)
Fine-tuning CLIP Text Encoders with Two-step Paraphrasing
by: Kim, Hyunjae, et al.
Published: (2024)
by: Kim, Hyunjae, et al.
Published: (2024)
Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIP
by: Eslami, Sedigheh, et al.
Published: (2024)
by: Eslami, Sedigheh, et al.
Published: (2024)
IDEA: Image Description Enhanced CLIP-Adapter
by: Ye, Zhipeng, et al.
Published: (2025)
by: Ye, Zhipeng, et al.
Published: (2025)
FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
by: Xie, Chunyu, et al.
Published: (2025)
by: Xie, Chunyu, et al.
Published: (2025)
Similar Items
-
Localizing Memorization in SSL Vision Encoders
by: Wang, Wenhao, et al.
Published: (2024) -
Finding DoRI: Discovery of Retained Images in Diffusion Models
by: Kowalczuk, Antoni, et al.
Published: (2025) -
Demystifying Foreground-Background Memorization in Diffusion Models
by: Di, Jimmy Z., et al.
Published: (2025) -
Privacy Attacks on Image AutoRegressive Models
by: Kowalczuk, Antoni, et al.
Published: (2025) -
Radioactive Watermarks in Diffusion and Autoregressive Image Generative Models
by: Meintz, Michel, et al.
Published: (2025)