Beyond Suffixes: Token Position in GCG Adversarial Attacks on Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Eddoubi, Hicham, Abdullahi, Umar Faruk, Hassan, Fadi
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913078914842624
author Eddoubi, Hicham
Abdullahi, Umar Faruk
Hassan, Fadi
author_facet Eddoubi, Hicham
Abdullahi, Umar Faruk
Hassan, Fadi
contents Large Language Models (LLMs) have seen widespread adoption across multiple domains, creating an urgent need for robust safety alignment mechanisms. However, robustness remains challenging due to jailbreak attacks that bypass alignment via adversarial prompts. In this work, we focus on the prevalent Greedy Coordinate Gradient (GCG) attack and identify a previously underexplored attack axis in jailbreak attacks typically framed as suffix-based: the placement of adversarial tokens within the prompt. Using GCG as a case study, we show that both optimizing attacks to generate prefixes instead of suffixes and varying adversarial token position during evaluation substantially influence attack success rates. Our findings highlight a critical blind spot in current safety evaluations and underline the need to account for the position of adversarial tokens in the adversarial robustness evaluation of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03265
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Suffixes: Token Position in GCG Adversarial Attacks on Large Language Models
Eddoubi, Hicham
Abdullahi, Umar Faruk
Hassan, Fadi
Machine Learning
Large Language Models (LLMs) have seen widespread adoption across multiple domains, creating an urgent need for robust safety alignment mechanisms. However, robustness remains challenging due to jailbreak attacks that bypass alignment via adversarial prompts. In this work, we focus on the prevalent Greedy Coordinate Gradient (GCG) attack and identify a previously underexplored attack axis in jailbreak attacks typically framed as suffix-based: the placement of adversarial tokens within the prompt. Using GCG as a case study, we show that both optimizing attacks to generate prefixes instead of suffixes and varying adversarial token position during evaluation substantially influence attack success rates. Our findings highlight a critical blind spot in current safety evaluations and underline the need to account for the position of adversarial tokens in the adversarial robustness evaluation of LLMs.
title Beyond Suffixes: Token Position in GCG Adversarial Attacks on Large Language Models
topic Machine Learning
url https://arxiv.org/abs/2602.03265