GCG Attack On A Diffusion LLM

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Neyroud, Ruben, Corley, Sam
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909996403392512
author Neyroud, Ruben
Corley, Sam
author_facet Neyroud, Ruben
Corley, Sam
contents While most LLMs are autoregressive, diffusion-based LLMs have recently emerged as an alternative method for generation. Greedy Coordinate Gradient (GCG) attacks have proven effective against autoregressive models, but their applicability to diffusion language models remains largely unexplored. In this work, we present an exploratory study of GCG-style adversarial prompt attacks on LLaDA (Large Language Diffusion with mAsking), an open-source diffusion LLM. We evaluate multiple attack variants, including prefix perturbations and suffix-based adversarial generation, on harmful prompts drawn from the AdvBench dataset. Our study provides initial insights into the robustness and attack surface of diffusion language models and motivates the development of alternative optimization and evaluation strategies for adversarial analysis in this setting.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14266
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GCG Attack On A Diffusion LLM
Neyroud, Ruben
Corley, Sam
Machine Learning
Computation and Language
Cryptography and Security
While most LLMs are autoregressive, diffusion-based LLMs have recently emerged as an alternative method for generation. Greedy Coordinate Gradient (GCG) attacks have proven effective against autoregressive models, but their applicability to diffusion language models remains largely unexplored. In this work, we present an exploratory study of GCG-style adversarial prompt attacks on LLaDA (Large Language Diffusion with mAsking), an open-source diffusion LLM. We evaluate multiple attack variants, including prefix perturbations and suffix-based adversarial generation, on harmful prompts drawn from the AdvBench dataset. Our study provides initial insights into the robustness and attack surface of diffusion language models and motivates the development of alternative optimization and evaluation strategies for adversarial analysis in this setting.
title GCG Attack On A Diffusion LLM
topic Machine Learning
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2601.14266