Guardado en:
| Autores principales: | Eddoubi, Hicham, Abdullahi, Umar Faruk, Hassan, Fadi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2602.03265 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Resurgence of GCG Adversarial Attacks on Large Language Models
por: Tan, Yuting, et al.
Publicado: (2025)
por: Tan, Yuting, et al.
Publicado: (2025)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
por: Mu, Junjie, et al.
Publicado: (2025)
por: Mu, Junjie, et al.
Publicado: (2025)
GCG Attack On A Diffusion LLM
por: Neyroud, Ruben, et al.
Publicado: (2025)
por: Neyroud, Ruben, et al.
Publicado: (2025)
Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models
por: Liu, Hongfu, et al.
Publicado: (2024)
por: Liu, Hongfu, et al.
Publicado: (2024)
Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models
por: Li, Xiao, et al.
Publicado: (2024)
por: Li, Xiao, et al.
Publicado: (2024)
RAID: A Dataset for Testing the Adversarial Robustness of AI-Generated Image Detectors
por: Eddoubi, Hicham, et al.
Publicado: (2025)
por: Eddoubi, Hicham, et al.
Publicado: (2025)
Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
por: Tang, Haochun, et al.
Publicado: (2026)
por: Tang, Haochun, et al.
Publicado: (2026)
Adversarial Suffix Filtering: a Defense Pipeline for LLMs
por: Khachaturov, David, et al.
Publicado: (2025)
por: Khachaturov, David, et al.
Publicado: (2025)
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
por: Yang, Xiaoxue, et al.
Publicado: (2025)
por: Yang, Xiaoxue, et al.
Publicado: (2025)
A Triadic Suffix Tokenization Scheme for Numerical Reasoning
por: Chetverina, Olga
Publicado: (2026)
por: Chetverina, Olga
Publicado: (2026)
BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks
por: Zhao, Yunhan, et al.
Publicado: (2024)
por: Zhao, Yunhan, et al.
Publicado: (2024)
Sampling-aware Adversarial Attacks Against Large Language Models
por: Beyer, Tim, et al.
Publicado: (2025)
por: Beyer, Tim, et al.
Publicado: (2025)
DPad: Efficient Diffusion Language Models with Suffix Dropout
por: Chen, Xinhua, et al.
Publicado: (2025)
por: Chen, Xinhua, et al.
Publicado: (2025)
Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language Models
por: Wang, Yubo, et al.
Publicado: (2024)
por: Wang, Yubo, et al.
Publicado: (2024)
A Closer Look at Adversarial Suffix Learning for Jailbreaking LLMs: Augmented Adversarial Trigger Learning
por: Wang, Zhe, et al.
Publicado: (2025)
por: Wang, Zhe, et al.
Publicado: (2025)
Adversarial Evasion Attack Efficiency against Large Language Models
por: Vitorino, João, et al.
Publicado: (2024)
por: Vitorino, João, et al.
Publicado: (2024)
Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
por: Roth, Tom, et al.
Publicado: (2021)
por: Roth, Tom, et al.
Publicado: (2021)
REINFORCE Adversarial Attacks on Large Language Models: An Adaptive, Distributional, and Semantic Objective
por: Geisler, Simon, et al.
Publicado: (2025)
por: Geisler, Simon, et al.
Publicado: (2025)
Beyond Semantic Manipulation: Token-Space Attacks on Reward Models
por: Zhang, Yuheng, et al.
Publicado: (2026)
por: Zhang, Yuheng, et al.
Publicado: (2026)
Beyond Early-Token Bias: Model-Specific and Language-Specific Position Effects in Multilingual LLMs
por: Menschikov, Mikhail, et al.
Publicado: (2025)
por: Menschikov, Mikhail, et al.
Publicado: (2025)
Evaluating the Performance of Large Language Models in Scientific Claim Detection and Classification
por: Faruk, Tanjim Bin
Publicado: (2024)
por: Faruk, Tanjim Bin
Publicado: (2024)
Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent
por: Biswas, Sajib, et al.
Publicado: (2025)
por: Biswas, Sajib, et al.
Publicado: (2025)
Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
por: Winninger, Thomas, et al.
Publicado: (2025)
por: Winninger, Thomas, et al.
Publicado: (2025)
Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
por: Biswas, Sajib, et al.
Publicado: (2025)
por: Biswas, Sajib, et al.
Publicado: (2025)
Policy Disruption in Reinforcement Learning:Adversarial Attack with Large Language Models and Critical State Identification
por: Jiang, Junyong, et al.
Publicado: (2025)
por: Jiang, Junyong, et al.
Publicado: (2025)
Tokens for Learning, Tokens for Unlearning: Mitigating Membership Inference Attacks in Large Language Models via Dual-Purpose Training
por: Tran, Toan, et al.
Publicado: (2025)
por: Tran, Toan, et al.
Publicado: (2025)
GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs
por: Basani, Advik Raj, et al.
Publicado: (2024)
por: Basani, Advik Raj, et al.
Publicado: (2024)
Detecting Dark Patterns in User Interfaces Using Logistic Regression and Bag-of-Words Representation
por: Umar, Aliyu, et al.
Publicado: (2024)
por: Umar, Aliyu, et al.
Publicado: (2024)
Vision Transformer with Adversarial Indicator Token against Adversarial Attacks in Radio Signal Classifications
por: Zhang, Lu, et al.
Publicado: (2025)
por: Zhang, Lu, et al.
Publicado: (2025)
Adversarial Attacks on Audio Deepfake Detection: A Benchmark and Comparative Study
por: Uddin, Kutub, et al.
Publicado: (2025)
por: Uddin, Kutub, et al.
Publicado: (2025)
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
por: Shao, Chenze, et al.
Publicado: (2024)
por: Shao, Chenze, et al.
Publicado: (2024)
INTERPOS: Interaction Rhythm Guided Positional Morphing for Mobile App Recommender Systems
por: Maqbool, M. H., et al.
Publicado: (2025)
por: Maqbool, M. H., et al.
Publicado: (2025)
Adversarial Attacks on Large Language Models Using Regularized Relaxation
por: Chacko, Samuel Jacob, et al.
Publicado: (2024)
por: Chacko, Samuel Jacob, et al.
Publicado: (2024)
DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models
por: Sun, Ye, et al.
Publicado: (2026)
por: Sun, Ye, et al.
Publicado: (2026)
Consistent Valid Physically-Realizable Adversarial Attack against Crowd-flow Prediction Models
por: Ali, Hassan, et al.
Publicado: (2023)
por: Ali, Hassan, et al.
Publicado: (2023)
Time-To-Inconsistency: A Survival Analysis of Large Language Model Robustness to Adversarial Attacks
por: Li, Yubo, et al.
Publicado: (2025)
por: Li, Yubo, et al.
Publicado: (2025)
AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs
por: Liao, Zeyi, et al.
Publicado: (2024)
por: Liao, Zeyi, et al.
Publicado: (2024)
AnyAttack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models
por: Zhang, Jiaming, et al.
Publicado: (2024)
por: Zhang, Jiaming, et al.
Publicado: (2024)
Revisiting Character-level Adversarial Attacks for Language Models
por: Rocamora, Elias Abad, et al.
Publicado: (2024)
por: Rocamora, Elias Abad, et al.
Publicado: (2024)
ER-MIA: Black-Box Adversarial Memory Injection Attacks on Long-Term Memory-Augmented Large Language Models
por: Piehl, Mitchell, et al.
Publicado: (2026)
por: Piehl, Mitchell, et al.
Publicado: (2026)
Ejemplares similares
-
The Resurgence of GCG Adversarial Attacks on Large Language Models
por: Tan, Yuting, et al.
Publicado: (2025) -
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
por: Mu, Junjie, et al.
Publicado: (2025) -
GCG Attack On A Diffusion LLM
por: Neyroud, Ruben, et al.
Publicado: (2025) -
Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models
por: Liu, Hongfu, et al.
Publicado: (2024) -
Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models
por: Li, Xiao, et al.
Publicado: (2024)