FameBias: Embedding Manipulation Bias Attack in Text-to-Image Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Roh, Jaechul, Yuan, Andrew, Mao, Jinsong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929646350630912
author Roh, Jaechul
Yuan, Andrew
Mao, Jinsong
author_facet Roh, Jaechul
Yuan, Andrew
Mao, Jinsong
contents Text-to-Image (T2I) diffusion models have rapidly advanced, enabling the generation of high-quality images that align closely with textual descriptions. However, this progress has also raised concerns about their misuse for propaganda and other malicious activities. Recent studies reveal that attackers can embed biases into these models through simple fine-tuning, causing them to generate targeted imagery when triggered by specific phrases. This underscores the potential for T2I models to act as tools for disseminating propaganda, producing images aligned with an attacker's objective for end-users. Building on this concept, we introduce FameBias, a T2I biasing attack that manipulates the embeddings of input prompts to generate images featuring specific public figures. Unlike prior methods, Famebias operates solely on the input embedding vectors without requiring additional model training. We evaluate FameBias comprehensively using Stable Diffusion V2, generating a large corpus of images based on various trigger nouns and target public figures. Our experiments demonstrate that FameBias achieves a high attack success rate while preserving the semantic context of the original prompts across multiple trigger-target pairs.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18302
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FameBias: Embedding Manipulation Bias Attack in Text-to-Image Models
Roh, Jaechul
Yuan, Andrew
Mao, Jinsong
Computer Vision and Pattern Recognition
Cryptography and Security
Machine Learning
Text-to-Image (T2I) diffusion models have rapidly advanced, enabling the generation of high-quality images that align closely with textual descriptions. However, this progress has also raised concerns about their misuse for propaganda and other malicious activities. Recent studies reveal that attackers can embed biases into these models through simple fine-tuning, causing them to generate targeted imagery when triggered by specific phrases. This underscores the potential for T2I models to act as tools for disseminating propaganda, producing images aligned with an attacker's objective for end-users. Building on this concept, we introduce FameBias, a T2I biasing attack that manipulates the embeddings of input prompts to generate images featuring specific public figures. Unlike prior methods, Famebias operates solely on the input embedding vectors without requiring additional model training. We evaluate FameBias comprehensively using Stable Diffusion V2, generating a large corpus of images based on various trigger nouns and target public figures. Our experiments demonstrate that FameBias achieves a high attack success rate while preserving the semantic context of the original prompts across multiple trigger-target pairs.
title FameBias: Embedding Manipulation Bias Attack in Text-to-Image Models
topic Computer Vision and Pattern Recognition
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2412.18302