Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cheng, Hao, Xiao, Erjia, Yang, Jiayan, Cao, Jiahang, Zhang, Qiang, Zhang, Jize, Xu, Kaidi, Gu, Jindong, Xu, Renjing
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909597357309952
author Cheng, Hao
Xiao, Erjia
Yang, Jiayan
Cao, Jiahang
Zhang, Qiang
Zhang, Jize
Xu, Kaidi
Gu, Jindong
Xu, Renjing
author_facet Cheng, Hao
Xiao, Erjia
Yang, Jiayan
Cao, Jiahang
Zhang, Qiang
Zhang, Jize
Xu, Kaidi
Gu, Jindong
Xu, Renjing
contents Current image generation models can effortlessly produce high-quality, highly realistic images, but this also increases the risk of misuse. In various Text-to-Image or Image-to-Image tasks, attackers can generate a series of images containing inappropriate content by simply editing the language modality input. To mitigate this security concern, numerous guarding or defensive strategies have been proposed, with a particular emphasis on safeguarding language modality. However, in practical applications, threats in the vision modality, particularly in tasks involving the editing of real-world images, present heightened security risks as they can easily infringe upon the rights of the image owner. Therefore, this paper employs a method named typographic attack to reveal that various image generation models are also susceptible to threats within the vision modality. Furthermore, we also evaluate the defense performance of various existing methods when facing threats in the vision modality and uncover their ineffectiveness. Finally, we propose the Vision Modal Threats in Image Generation Models (VMT-IGMs) dataset, which would serve as a baseline for evaluating the vision modality vulnerability of various image generation models.
format Preprint
id arxiv_https___arxiv_org_abs_2412_05538
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models
Cheng, Hao
Xiao, Erjia
Yang, Jiayan
Cao, Jiahang
Zhang, Qiang
Zhang, Jize
Xu, Kaidi
Gu, Jindong
Xu, Renjing
Computer Vision and Pattern Recognition
Performance
Current image generation models can effortlessly produce high-quality, highly realistic images, but this also increases the risk of misuse. In various Text-to-Image or Image-to-Image tasks, attackers can generate a series of images containing inappropriate content by simply editing the language modality input. To mitigate this security concern, numerous guarding or defensive strategies have been proposed, with a particular emphasis on safeguarding language modality. However, in practical applications, threats in the vision modality, particularly in tasks involving the editing of real-world images, present heightened security risks as they can easily infringe upon the rights of the image owner. Therefore, this paper employs a method named typographic attack to reveal that various image generation models are also susceptible to threats within the vision modality. Furthermore, we also evaluate the defense performance of various existing methods when facing threats in the vision modality and uncover their ineffectiveness. Finally, we propose the Vision Modal Threats in Image Generation Models (VMT-IGMs) dataset, which would serve as a baseline for evaluating the vision modality vulnerability of various image generation models.
title Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models
topic Computer Vision and Pattern Recognition
Performance
url https://arxiv.org/abs/2412.05538