Semantic-Preserving Adversarial Attacks on LLMs: An Adaptive Greedy Binary Search Approach
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Chong, Li, Xiang, Wang, Jia, Liang, Shan, Xue, Haochen, Jin, Xiaobo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Goal-guided Generative Prompt Injection Attack on Large Language Models
by: Zhang, Chong, et al.
Published: (2024)
by: Zhang, Chong, et al.
Published: (2024)
AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
by: Lv, Lijia, et al.
Published: (2024)
by: Lv, Lijia, et al.
Published: (2024)
Shadow Unlearning: A Neuro-Semantic Approach to Fidelity-Preserving Faceless Forgetting in LLMs
by: P, Dinesh Srivasthav, et al.
Published: (2026)
by: P, Dinesh Srivasthav, et al.
Published: (2026)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
by: Chen, Yunhao, et al.
Published: (2025)
by: Chen, Yunhao, et al.
Published: (2025)
Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
by: Fastowski, Alina, et al.
Published: (2025)
by: Fastowski, Alina, et al.
Published: (2025)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
by: Brown, Hannah, et al.
Published: (2024)
by: Brown, Hannah, et al.
Published: (2024)
Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training
by: Li, Yuanfan, et al.
Published: (2025)
by: Li, Yuanfan, et al.
Published: (2025)
Semantic Stealth: Adversarial Text Attacks on NLP Using Several Methods
by: Dey, Roopkatha, et al.
Published: (2024)
by: Dey, Roopkatha, et al.
Published: (2024)
BadApex: Backdoor Attack Based on Adaptive Optimization Mechanism of Black-box Large Language Models
by: Wu, Zhengxian, et al.
Published: (2025)
by: Wu, Zhengxian, et al.
Published: (2025)
Enhancing Adversarial Attacks via Parameter Adaptive Adversarial Attack
by: Jin, Zhibo, et al.
Published: (2024)
by: Jin, Zhibo, et al.
Published: (2024)
BinarySelect to Improve Accessibility of Black-Box Attack Research
by: Ghosh, Shatarupa, et al.
Published: (2024)
by: Ghosh, Shatarupa, et al.
Published: (2024)
Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs
by: Upadhayay, Bibek, et al.
Published: (2024)
by: Upadhayay, Bibek, et al.
Published: (2024)
MPAT: Building Robust Deep Neural Networks against Textual Adversarial Attacks
by: Zhang, Fangyuan, et al.
Published: (2024)
by: Zhang, Fangyuan, et al.
Published: (2024)
Imposter.AI: Adversarial Attacks with Hidden Intentions towards Aligned Large Language Models
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
by: Kaneko, Masahiro, et al.
Published: (2025)
by: Kaneko, Masahiro, et al.
Published: (2025)
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs
by: Berezin, Sergey, et al.
Published: (2025)
by: Berezin, Sergey, et al.
Published: (2025)
Emoti-Attack: Zero-Perturbation Adversarial Attacks on NLP Systems via Emoji Sequences
by: Zhang, Yangshijie
Published: (2025)
by: Zhang, Yangshijie
Published: (2025)
Look Twice before You Leap: A Rational Framework for Localized Adversarial Anonymization
by: Duan, Donghang, et al.
Published: (2025)
by: Duan, Donghang, et al.
Published: (2025)
Continual Pretraining on Encrypted Synthetic Data for Privacy-Preserving LLMs
by: Liu, Honghao, et al.
Published: (2026)
by: Liu, Honghao, et al.
Published: (2026)
Adversarial Malware Generation in Linux ELF Binaries via Semantic-Preserving Transformations
by: Hrdonka, Lukáš, et al.
Published: (2026)
by: Hrdonka, Lukáš, et al.
Published: (2026)
From Essence to Defense: Adaptive Semantic-aware Watermarking for Embedding-as-a-Service Copyright Protection
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
by: Maloyan, Narek, et al.
Published: (2025)
by: Maloyan, Narek, et al.
Published: (2025)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Adversarial Attacks Against Automated Fact-Checking: A Survey
by: Liu, Fanzhen, et al.
Published: (2025)
by: Liu, Fanzhen, et al.
Published: (2025)
SEEP: Training Dynamics Grounds Latent Representation Search for Mitigating Backdoor Poisoning Attacks
by: He, Xuanli, et al.
Published: (2024)
by: He, Xuanli, et al.
Published: (2024)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
by: Wang, Jiawen, et al.
Published: (2025)
by: Wang, Jiawen, et al.
Published: (2025)
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
by: Wang, Peiran, et al.
Published: (2026)
by: Wang, Peiran, et al.
Published: (2026)
TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning
by: He, Xuanli, et al.
Published: (2024)
by: He, Xuanli, et al.
Published: (2024)
Multi-Granularity Tibetan Textual Adversarial Attack Method Based on Masked Language Model
by: Cao, Xi, et al.
Published: (2024)
by: Cao, Xi, et al.
Published: (2024)
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
by: Wang, Yidan, et al.
Published: (2025)
by: Wang, Yidan, et al.
Published: (2025)
LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation
by: Li, Xingyu, et al.
Published: (2025)
by: Li, Xingyu, et al.
Published: (2025)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
by: Shen, Guobin, et al.
Published: (2025)
by: Shen, Guobin, et al.
Published: (2025)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
by: Jia, Yuqi, et al.
Published: (2026)
by: Jia, Yuqi, et al.
Published: (2026)
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
by: Wei, Zhang, et al.
Published: (2025)
by: Wei, Zhang, et al.
Published: (2025)
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
by: Li, Ran, et al.
Published: (2025)
by: Li, Ran, et al.
Published: (2025)
A Content-Preserving Secure Linguistic Steganography
by: Xiang, Lingyun, et al.
Published: (2025)
by: Xiang, Lingyun, et al.
Published: (2025)
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
by: Paulus, Anselm, et al.
Published: (2024)
by: Paulus, Anselm, et al.
Published: (2024)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
by: Mu, Junjie, et al.
Published: (2025)
by: Mu, Junjie, et al.
Published: (2025)
Similar Items
-
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
by: Li, Xiang, et al.
Published: (2025) -
Goal-guided Generative Prompt Injection Attack on Large Language Models
by: Zhang, Chong, et al.
Published: (2024) -
AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
by: Lv, Lijia, et al.
Published: (2024) -
Shadow Unlearning: A Neuro-Semantic Approach to Fidelity-Preserving Faceless Forgetting in LLMs
by: P, Dinesh Srivasthav, et al.
Published: (2026) -
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)