An indicator for effectiveness of text-to-image guardrails utilizing the Single-Turn Crescendo Attack (STCA)
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915039064096768 |
|---|---|
| author | Kwartler, Ted Bagan, Nataliia Banny, Ivan Aqrawi, Alan Abbasi, Arian |
| author_facet | Kwartler, Ted Bagan, Nataliia Banny, Ivan Aqrawi, Alan Abbasi, Arian |
| contents | The Single-Turn Crescendo Attack (STCA), first introduced in Aqrawi and Abbasi [2024], is an innovative method designed to bypass the ethical safeguards of text-to-text AI models, compelling them to generate harmful content. This technique leverages a strategic escalation of context within a single prompt, combined with trust-building mechanisms, to subtly deceive the model into producing unintended outputs. Extending the application of STCA to text-to-image models, we demonstrate its efficacy by compromising the guardrails of a widely-used model, DALL-E 3, achieving outputs comparable to outputs from the uncensored model Flux Schnell, which served as a baseline control. This study provides a framework for researchers to rigorously evaluate the robustness of guardrails in text-to-image models and benchmark their resilience against adversarial attacks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_18699 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | An indicator for effectiveness of text-to-image guardrails utilizing the Single-Turn Crescendo Attack (STCA) Kwartler, Ted Bagan, Nataliia Banny, Ivan Aqrawi, Alan Abbasi, Arian Cryptography and Security Computation and Language The Single-Turn Crescendo Attack (STCA), first introduced in Aqrawi and Abbasi [2024], is an innovative method designed to bypass the ethical safeguards of text-to-text AI models, compelling them to generate harmful content. This technique leverages a strategic escalation of context within a single prompt, combined with trust-building mechanisms, to subtly deceive the model into producing unintended outputs. Extending the application of STCA to text-to-image models, we demonstrate its efficacy by compromising the guardrails of a widely-used model, DALL-E 3, achieving outputs comparable to outputs from the uncensored model Flux Schnell, which served as a baseline control. This study provides a framework for researchers to rigorously evaluate the robustness of guardrails in text-to-image models and benchmark their resilience against adversarial attacks. |
| title | An indicator for effectiveness of text-to-image guardrails utilizing the Single-Turn Crescendo Attack (STCA) |
| topic | Cryptography and Security Computation and Language |
| url | https://arxiv.org/abs/2411.18699 |