DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Ye, Wang, Xin, Zhang, Jiaming, Gao, Yifeng, Wang, Yixu, Ding, Yifan, Zhang, Qixian, Ding, Henghui, Ma, Xingjun, Jiang, Yu-Gang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data
by: Wang, Yixu, et al.
Published: (2025)
by: Wang, Yixu, et al.
Published: (2025)
Large Language Model Adversarial Landscape Through the Lens of Attack Objectives
by: Wang, Nan, et al.
Published: (2025)
by: Wang, Nan, et al.
Published: (2025)
ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
by: Wu, Yutao, et al.
Published: (2025)
by: Wu, Yutao, et al.
Published: (2025)
HoneypotNet: Backdoor Attacks Against Model Extraction
by: Wang, Yixu, et al.
Published: (2025)
by: Wang, Yixu, et al.
Published: (2025)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
by: Chen, Yunhao, et al.
Published: (2025)
by: Chen, Yunhao, et al.
Published: (2025)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
by: Li, Jiayu, et al.
Published: (2025)
by: Li, Jiayu, et al.
Published: (2025)
EnJa: Ensemble Jailbreak on Large Language Models
by: Zhang, Jiahao, et al.
Published: (2024)
by: Zhang, Jiahao, et al.
Published: (2024)
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
by: Wang, Zilong, et al.
Published: (2025)
by: Wang, Zilong, et al.
Published: (2025)
Internal Safety Collapse in Frontier Large Language Models
by: Wu, Yutao, et al.
Published: (2026)
by: Wu, Yutao, et al.
Published: (2026)
The Gradient Puppeteer: Adversarial Domination in Gradient Leakage Attacks through Model Poisoning
by: Xiang, Kunlan, et al.
Published: (2025)
by: Xiang, Kunlan, et al.
Published: (2025)
Membership Inference Attacks Against Video Large Language Models
by: Song, Wei, et al.
Published: (2026)
by: Song, Wei, et al.
Published: (2026)
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
by: Xu, Zonghuan, et al.
Published: (2025)
by: Xu, Zonghuan, et al.
Published: (2025)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
by: Zhao, Yunhan, et al.
Published: (2026)
by: Zhao, Yunhan, et al.
Published: (2026)
Transferable Adversarial Attacks on SAM and Its Downstream Models
by: Xia, Song, et al.
Published: (2024)
by: Xia, Song, et al.
Published: (2024)
LeakyCLIP: Extracting Training Data from CLIP
by: Chen, Yunhao, et al.
Published: (2025)
by: Chen, Yunhao, et al.
Published: (2025)
BadTemplate: A Training-Free Backdoor Attack via Chat Template Against Large Language Models
by: Wang, Zihan, et al.
Published: (2026)
by: Wang, Zihan, et al.
Published: (2026)
Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense
by: Hao, Shuyang, et al.
Published: (2025)
by: Hao, Shuyang, et al.
Published: (2025)
On Large Language Model Continual Unlearning
by: Gao, Chongyang, et al.
Published: (2024)
by: Gao, Chongyang, et al.
Published: (2024)
Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical Evidence
by: Fu, Shaopeng, et al.
Published: (2025)
by: Fu, Shaopeng, et al.
Published: (2025)
Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models
by: Ni, Zhenyang, et al.
Published: (2024)
by: Ni, Zhenyang, et al.
Published: (2024)
Unique Security and Privacy Threats of Large Language Models: A Comprehensive Survey
by: Wang, Shang, et al.
Published: (2024)
by: Wang, Shang, et al.
Published: (2024)
Imperceptible Jailbreaking against Large Language Models
by: Gao, Kuofeng, et al.
Published: (2025)
by: Gao, Kuofeng, et al.
Published: (2025)
Infighting in the Dark: Multi-Label Backdoor Attack in Federated Learning
by: Li, Ye, et al.
Published: (2024)
by: Li, Ye, et al.
Published: (2024)
Counterfactual Explainable Incremental Prompt Attack Analysis on Large Language Models
by: Shu, Dong, et al.
Published: (2024)
by: Shu, Dong, et al.
Published: (2024)
AttackEval: A Systematic Empirical Study of Prompt Injection Attack Effectiveness Against Large Language Models
by: Wang, Jackson
Published: (2026)
by: Wang, Jackson
Published: (2026)
Towards Action Hijacking of Large Language Model-based Agent
by: Zhang, Yuyang, et al.
Published: (2024)
by: Zhang, Yuyang, et al.
Published: (2024)
T2UE: Generating Unlearnable Examples from Text Descriptions
by: Ma, Xingjun, et al.
Published: (2025)
by: Ma, Xingjun, et al.
Published: (2025)
The Dark Side of Function Calling: Pathways to Jailbreaking Large Language Models
by: Wu, Zihui, et al.
Published: (2024)
by: Wu, Zihui, et al.
Published: (2024)
Adversarial Attacks on Multimodal Large Language Models: A Comprehensive Survey
by: Jain, Bhavuk, et al.
Published: (2026)
by: Jain, Bhavuk, et al.
Published: (2026)
Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language Models
by: He, Jiaming, et al.
Published: (2024)
by: He, Jiaming, et al.
Published: (2024)
Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models
by: Chen, Zeyuan, et al.
Published: (2026)
by: Chen, Zeyuan, et al.
Published: (2026)
Attack by Unlearning: Unlearning-Induced Adversarial Attacks on Graph Neural Networks
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
DEFENDCLI: {Command-Line} Driven Attack Provenance Examination
by: Wu, Peilun, et al.
Published: (2025)
by: Wu, Peilun, et al.
Published: (2025)
HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
by: Liu, Yuexiao, et al.
Published: (2025)
by: Liu, Yuexiao, et al.
Published: (2025)
SwitchPatch: Physical Adversarial Attack Strategy with Switchable Adversarial Objectives
by: Jiang, Hanrui, et al.
Published: (2025)
by: Jiang, Hanrui, et al.
Published: (2025)
Shortcuts Everywhere and Nowhere: Exploring Multi-Trigger Backdoor Attacks
by: Li, Yige, et al.
Published: (2024)
by: Li, Yige, et al.
Published: (2024)
LLM-Driven Feature-Level Adversarial Attacks on Android Malware Detectors
by: Lan, Tianwei, et al.
Published: (2025)
by: Lan, Tianwei, et al.
Published: (2025)
JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation
by: Zhang, Shenyi, et al.
Published: (2025)
by: Zhang, Shenyi, et al.
Published: (2025)
Towards Robust Multimodal Large Language Models Against Jailbreak Attacks
by: Yin, Ziyi, et al.
Published: (2025)
by: Yin, Ziyi, et al.
Published: (2025)
Similar Items
-
StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data
by: Wang, Yixu, et al.
Published: (2025) -
Large Language Model Adversarial Landscape Through the Lens of Attack Objectives
by: Wang, Nan, et al.
Published: (2025) -
ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
by: Wu, Yutao, et al.
Published: (2025) -
HoneypotNet: Backdoor Attacks Against Model Extraction
by: Wang, Yixu, et al.
Published: (2025) -
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
by: Chen, Yunhao, et al.
Published: (2025)