BadApex: Backdoor Attack Based on Adaptive Optimization Mechanism of Black-box Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Zhengxian, Wen, Juan, Peng, Wanli, Zhang, Ziwei, Zhou, Yinghan, Xue, Yiming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIs
by: Wu, Zhengxian, et al.
Published: (2025)
by: Wu, Zhengxian, et al.
Published: (2025)
BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge Distillation
by: Wu, Zhengxian, et al.
Published: (2025)
by: Wu, Zhengxian, et al.
Published: (2025)
Inhibitory Attacks on Backdoor-based Fingerprinting for Large Language Models
by: Fu, Hang, et al.
Published: (2026)
by: Fu, Hang, et al.
Published: (2026)
Self-Disguise Attack: Induce the LLM to disguise itself for AIGT detection evasion
by: Zhou, Yinghan, et al.
Published: (2025)
by: Zhou, Yinghan, et al.
Published: (2025)
GTSD: Generative Text Steganography Based on Diffusion Model
by: Wu, Zhengxian, et al.
Published: (2025)
by: Wu, Zhengxian, et al.
Published: (2025)
Is Your Writing Being Mimicked by AI? Unveiling Imitation with Invisible Watermarks in Creative Writing
by: Zhang, Ziwei, et al.
Published: (2025)
by: Zhang, Ziwei, et al.
Published: (2025)
EditMF: Drawing an Invisible Fingerprint for Your Large Language Models
by: Wu, Jiaxuan, et al.
Published: (2025)
by: Wu, Jiaxuan, et al.
Published: (2025)
BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models
by: Yuan, Zenghui, et al.
Published: (2025)
by: Yuan, Zenghui, et al.
Published: (2025)
Retrieval-Confused Generation is a Good Defender for Privacy Violation Attack of Large Language Models
by: Peng, Wanli, et al.
Published: (2025)
by: Peng, Wanli, et al.
Published: (2025)
Dynamic Black-box Backdoor Attacks on IoT Sensory Data
by: Chathoth, Ajesh Koyatan, et al.
Published: (2025)
by: Chathoth, Ajesh Koyatan, et al.
Published: (2025)
BadTemplate: A Training-Free Backdoor Attack via Chat Template Against Large Language Models
by: Wang, Zihan, et al.
Published: (2026)
by: Wang, Zihan, et al.
Published: (2026)
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
by: Yi, Biao, et al.
Published: (2025)
by: Yi, Biao, et al.
Published: (2025)
Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models
by: Chen, Zeyuan, et al.
Published: (2026)
by: Chen, Zeyuan, et al.
Published: (2026)
BadPatches: Routing-aware Backdoor Attacks on Vision Mixture of Experts
by: Chan, Cedric, et al.
Published: (2025)
by: Chan, Cedric, et al.
Published: (2025)
MalModel: Hiding Malicious Payload in Mobile Deep Learning Models with Black-box Backdoor Attack
by: Hua, Jiayi, et al.
Published: (2024)
by: Hua, Jiayi, et al.
Published: (2024)
AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
by: Li, Yanjie, et al.
Published: (2025)
by: Li, Yanjie, et al.
Published: (2025)
BadMerging: Backdoor Attacks Against Model Merging
by: Zhang, Jinghuai, et al.
Published: (2024)
by: Zhang, Jinghuai, et al.
Published: (2024)
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
by: Tie, Guiyao, et al.
Published: (2026)
by: Tie, Guiyao, et al.
Published: (2026)
BadTime: An Effective Backdoor Attack on Multivariate Long-Term Time Series Forecasting
by: Xiang, Kunlan, et al.
Published: (2025)
by: Xiang, Kunlan, et al.
Published: (2025)
BadDLM: Backdooring Diffusion Language Models with Diverse Targets
by: Zhai, Shengfang, et al.
Published: (2026)
by: Zhai, Shengfang, et al.
Published: (2026)
Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models
by: Yuan, Zenghui, et al.
Published: (2025)
by: Yuan, Zenghui, et al.
Published: (2025)
BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers
by: Xue, Jiaqi, et al.
Published: (2024)
by: Xue, Jiaqi, et al.
Published: (2024)
TooBadRL: Trigger Optimization to Boost Effectiveness of Backdoor Attacks on Deep Reinforcement Learning
by: Zhang, Mingxuan, et al.
Published: (2025)
by: Zhang, Mingxuan, et al.
Published: (2025)
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
by: Xiang, Zhen, et al.
Published: (2024)
by: Xiang, Zhen, et al.
Published: (2024)
Backdoors in Code Summarizers: How Bad Is It?
by: Wang, Chenyu, et al.
Published: (2025)
by: Wang, Chenyu, et al.
Published: (2025)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
BadRSSD: Backdoor Attacks on Regularized Self-Supervised Diffusion Models
by: Wang, Jiayao, et al.
Published: (2026)
by: Wang, Jiayao, et al.
Published: (2026)
BadActs: A Universal Backdoor Defense in the Activation Space
by: Yi, Biao, et al.
Published: (2024)
by: Yi, Biao, et al.
Published: (2024)
BadDet+: Robust Backdoor Attacks for Object Detection
by: Dunnett, Kealan, et al.
Published: (2026)
by: Dunnett, Kealan, et al.
Published: (2026)
BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning
by: Yang, Ziqing, et al.
Published: (2026)
by: Yang, Ziqing, et al.
Published: (2026)
Isolate Trigger: Detecting and Eliminating Adaptive Backdoor Attacks
by: Sun, Chengrui, et al.
Published: (2025)
by: Sun, Chengrui, et al.
Published: (2025)
When Reasoning Leaks Membership: Membership Inference Attack on Black-box Large Reasoning Models
by: Hu, Ruihan, et al.
Published: (2026)
by: Hu, Ruihan, et al.
Published: (2026)
BadSNN: Backdoor Attacks on Spiking Neural Networks via Adversarial Spiking Neuron
by: Miah, Abdullah Arafat, et al.
Published: (2026)
by: Miah, Abdullah Arafat, et al.
Published: (2026)
Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models
by: Ni, Zhenyang, et al.
Published: (2024)
by: Ni, Zhenyang, et al.
Published: (2024)
Double Backdoored: Converting Code Large Language Model Backdoors to Traditional Malware via Adversarial Instruction Tuning Attacks
by: Hossen, Md Imran, et al.
Published: (2024)
by: Hossen, Md Imran, et al.
Published: (2024)
BlackboxBench: A Comprehensive Benchmark of Black-box Adversarial Attacks
by: Zheng, Meixi, et al.
Published: (2023)
by: Zheng, Meixi, et al.
Published: (2023)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
by: Xue, Eric, et al.
Published: (2025)
by: Xue, Eric, et al.
Published: (2025)
BadImplant: Injection-based Multi-Targeted Graph Backdoor Attack
by: Khan, Md Nabi Newaz, et al.
Published: (2026)
by: Khan, Md Nabi Newaz, et al.
Published: (2026)
When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations
by: Ge, Huaizhi, et al.
Published: (2024)
by: Ge, Huaizhi, et al.
Published: (2024)
Similar Items
-
SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIs
by: Wu, Zhengxian, et al.
Published: (2025) -
BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge Distillation
by: Wu, Zhengxian, et al.
Published: (2025) -
Inhibitory Attacks on Backdoor-based Fingerprinting for Large Language Models
by: Fu, Hang, et al.
Published: (2026) -
Self-Disguise Attack: Induce the LLM to disguise itself for AIGT detection evasion
by: Zhou, Yinghan, et al.
Published: (2025) -
GTSD: Generative Text Steganography Based on Diffusion Model
by: Wu, Zhengxian, et al.
Published: (2025)