Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Kuo, Kevin, Yadav, Chhavi, Smith, Virginia |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attacks and Defenses Against LLM Fingerprinting
by: Kurian, Kevin, et al.
Published: (2025)
by: Kurian, Kevin, et al.
Published: (2025)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
by: Zhan, Qiusi, et al.
Published: (2025)
by: Zhan, Qiusi, et al.
Published: (2025)
Dummy-Aware Weighted Attack (DAWA): Breaking the Safe Sink in Dummy Class Defenses
by: Yu, Yunrui, et al.
Published: (2026)
by: Yu, Yunrui, et al.
Published: (2026)
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
by: Hossain, S M Asif, et al.
Published: (2025)
by: Hossain, S M Asif, et al.
Published: (2025)
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
by: Yang, Xiaoxue, et al.
Published: (2025)
by: Yang, Xiaoxue, et al.
Published: (2025)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024)
by: Zeng, Yifan, et al.
Published: (2024)
Can We Infer Confidential Properties of Training Data from LLMs?
by: Huang, Pengrun, et al.
Published: (2025)
by: Huang, Pengrun, et al.
Published: (2025)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
by: Debenedetti, Edoardo, et al.
Published: (2024)
by: Debenedetti, Edoardo, et al.
Published: (2024)
Membership Inference Attacks for Unseen Classes
by: Thaker, Pratiksha, et al.
Published: (2025)
by: Thaker, Pratiksha, et al.
Published: (2025)
PAE MobiLLM: Privacy-Aware and Efficient LLM Fine-Tuning on the Mobile Device via Additive Side-Tuning
by: Yang, Xingke, et al.
Published: (2025)
by: Yang, Xingke, et al.
Published: (2025)
Attack and Defense of Deep Learning Models in the Field of Web Attack Detection
by: Shi, Lijia, et al.
Published: (2024)
by: Shi, Lijia, et al.
Published: (2024)
FairProof : Confidential and Certifiable Fairness for Neural Networks
by: Yadav, Chhavi, et al.
Published: (2024)
by: Yadav, Chhavi, et al.
Published: (2024)
ExpProof : Operationalizing Explanations for Confidential Models with ZKPs
by: Yadav, Chhavi, et al.
Published: (2025)
by: Yadav, Chhavi, et al.
Published: (2025)
Synthetic Tabular Data: Methods, Attacks and Defenses
by: Cormode, Graham, et al.
Published: (2025)
by: Cormode, Graham, et al.
Published: (2025)
A Taxonomy of Attacks and Defenses in Split Learning
by: Shabbir, Aqsa, et al.
Published: (2025)
by: Shabbir, Aqsa, et al.
Published: (2025)
Early Approaches to Adversarial Fine-Tuning for Prompt Injection Defense: A 2022 Study of GPT-3 and Contemporary Models
by: Sandoval, Gustavo, et al.
Published: (2025)
by: Sandoval, Gustavo, et al.
Published: (2025)
MarkTune: Improving the Quality-Detectability Trade-off in Open-Weight LLM Watermarking
by: Zhao, Yizhou, et al.
Published: (2025)
by: Zhao, Yizhou, et al.
Published: (2025)
Dashed Line Defense: Plug-And-Play Defense Against Adaptive Score-Based Query Attacks
by: Fu, Yanzhang, et al.
Published: (2026)
by: Fu, Yanzhang, et al.
Published: (2026)
A No-Defense Defense Against Gradient-Based Adversarial Attacks on ML-NIDS: Is Less More?
by: elShehaby, Mohamed, et al.
Published: (2026)
by: elShehaby, Mohamed, et al.
Published: (2026)
Data Reconstruction Attacks and Defenses: A Systematic Evaluation
by: Liu, Sheng, et al.
Published: (2024)
by: Liu, Sheng, et al.
Published: (2024)
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework
by: Chu, Kexin
Published: (2026)
by: Chu, Kexin
Published: (2026)
Watch Out! Simple Horizontal Class Backdoor Can Trivially Evade Defense
by: Ma, Hua, et al.
Published: (2023)
by: Ma, Hua, et al.
Published: (2023)
Uncovering Attacks and Defenses in Secure Aggregation for Federated Deep Learning
by: Zhang, Yiwei, et al.
Published: (2024)
by: Zhang, Yiwei, et al.
Published: (2024)
PECAN: A Deterministic Certified Defense Against Backdoor Attacks
by: Zhang, Yuhao, et al.
Published: (2023)
by: Zhang, Yuhao, et al.
Published: (2023)
SecureLearn -- An Attack-agnostic Defense for Multiclass Machine Learning Against Data Poisoning Attacks
by: Paracha, Anum, et al.
Published: (2025)
by: Paracha, Anum, et al.
Published: (2025)
Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
by: Konrad, Phongsakon Mark, et al.
Published: (2026)
by: Konrad, Phongsakon Mark, et al.
Published: (2026)
Reconstruction Attacks on Machine Unlearning: Simple Models are Vulnerable
by: Bertran, Martin, et al.
Published: (2024)
by: Bertran, Martin, et al.
Published: (2024)
Exploring Secure Machine Learning Through Payload Injection and FGSM Attacks on ResNet-50
by: Yadav, Umesh, et al.
Published: (2025)
by: Yadav, Umesh, et al.
Published: (2025)
Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions
by: Li, Wenjuan, et al.
Published: (2026)
by: Li, Wenjuan, et al.
Published: (2026)
Inferring Sensitive Attributes from Knowledge Graph Embeddings: Attack and Defense Strategies
by: Hayder, Yasmine
Published: (2026)
by: Hayder, Yasmine
Published: (2026)
Attack Smarter: Attention-Driven Fine-Grained Webpage Fingerprinting Attacks
by: Yuan, Yali, et al.
Published: (2025)
by: Yuan, Yali, et al.
Published: (2025)
Exploiting Defenses against GAN-Based Feature Inference Attacks in Federated Learning
by: Luo, Xinjian, et al.
Published: (2020)
by: Luo, Xinjian, et al.
Published: (2020)
AttackLLM: LLM-based Attack Pattern Generation for an Industrial Control System
by: Ahmed, Chuadhry Mujeeb
Published: (2025)
by: Ahmed, Chuadhry Mujeeb
Published: (2025)
Fine-tuning is Not Fine: Mitigating Backdoor Attacks in GNNs with Limited Clean Data
by: Zhang, Jiale, et al.
Published: (2025)
by: Zhang, Jiale, et al.
Published: (2025)
SVDefense: Effective Defense against Gradient Inversion Attacks via Singular Value Decomposition
by: Luo, Chenxiang, et al.
Published: (2025)
by: Luo, Chenxiang, et al.
Published: (2025)
Test-time Adversarial Defense with Opposite Adversarial Path and High Attack Time Cost
by: Yeh, Cheng-Han, et al.
Published: (2024)
by: Yeh, Cheng-Han, et al.
Published: (2024)
Sponge Attacks on Sensing AI: Energy-Latency Vulnerabilities and Defense via Model Pruning
by: Hasan, Syed Mhamudul, et al.
Published: (2025)
by: Hasan, Syed Mhamudul, et al.
Published: (2025)
KDk: A Defense Mechanism Against Label Inference Attacks in Vertical Federated Learning
by: Arazzi, Marco, et al.
Published: (2024)
by: Arazzi, Marco, et al.
Published: (2024)
Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable Confidence
by: Hong, Hanbin, et al.
Published: (2023)
by: Hong, Hanbin, et al.
Published: (2023)
Similar Items
-
Attacks and Defenses Against LLM Fingerprinting
by: Kurian, Kevin, et al.
Published: (2025) -
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
by: Zhan, Qiusi, et al.
Published: (2025) -
Dummy-Aware Weighted Attack (DAWA): Breaking the Safe Sink in Dummy Class Defenses
by: Yu, Yunrui, et al.
Published: (2026) -
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
by: Hossain, S M Asif, et al.
Published: (2025) -
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
by: Yang, Xiaoxue, et al.
Published: (2025)