PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Zhen, Cong, Tianshuo, Liu, Yule, Lin, Chenhao, He, Xinlei, Chen, Rongmao, Han, Xingshuo, Huang, Xinyi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
by: Zheng, Jingyi, et al.
Published: (2024)
by: Zheng, Jingyi, et al.
Published: (2024)
Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models
by: Zhang, Zongmin, et al.
Published: (2025)
by: Zhang, Zongmin, et al.
Published: (2025)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
by: Zhang, Heyi, et al.
Published: (2025)
by: Zhang, Heyi, et al.
Published: (2025)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
by: Yi, Sibo, et al.
Published: (2024)
by: Yi, Sibo, et al.
Published: (2024)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
LoRA-Leak: Membership Inference Attacks Against LoRA Fine-tuned Language Models
by: Ran, Delong, et al.
Published: (2025)
by: Ran, Delong, et al.
Published: (2025)
Quantized Delta Weight Is Safety Keeper
by: Liu, Yule, et al.
Published: (2024)
by: Liu, Yule, et al.
Published: (2024)
Real is not True: Backdoor Attacks Against Deepfake Detection
by: Sun, Hong, et al.
Published: (2024)
by: Sun, Hong, et al.
Published: (2024)
On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
by: Liu, Zesen, et al.
Published: (2024)
by: Liu, Zesen, et al.
Published: (2024)
TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
by: Zheng, Jingyi, et al.
Published: (2025)
by: Zheng, Jingyi, et al.
Published: (2025)
Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation
by: Zhong, Zhiyuan, et al.
Published: (2025)
by: Zhong, Zhiyuan, et al.
Published: (2025)
Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs
by: Cui, Jing, et al.
Published: (2025)
by: Cui, Jing, et al.
Published: (2025)
Beyond the Tip of Efficiency: Uncovering the Submerged Threats of Jailbreak Attacks in Small Language Models
by: Yi, Sibo, et al.
Published: (2025)
by: Yi, Sibo, et al.
Published: (2025)
Auditing Data Membership in Reinforcement Learning With Verifiable Rewards
by: Liu, Yule, et al.
Published: (2025)
by: Liu, Yule, et al.
Published: (2025)
SSD: A State-based Stealthy Backdoor Attack For Navigation System in UAV Route Planning
by: Wang, Zhaoxuan, et al.
Published: (2025)
by: Wang, Zhaoxuan, et al.
Published: (2025)
Krait: A Backdoor Attack Against Graph Prompt Tuning
by: Song, Ying, et al.
Published: (2024)
by: Song, Ying, et al.
Published: (2024)
BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning
by: Yang, Ziqing, et al.
Published: (2026)
by: Yang, Ziqing, et al.
Published: (2026)
SwitchPatch: Physical Adversarial Attack Strategy with Switchable Adversarial Objectives
by: Jiang, Hanrui, et al.
Published: (2025)
by: Jiang, Hanrui, et al.
Published: (2025)
Activation Gradient based Poisoned Sample Detection Against Backdoor Attacks
by: Yuan, Danni, et al.
Published: (2023)
by: Yuan, Danni, et al.
Published: (2023)
Prediction Inconsistency Helps Achieve Generalizable Detection of Adversarial Examples
by: Han, Sicong, et al.
Published: (2025)
by: Han, Sicong, et al.
Published: (2025)
Meme Trojan: Backdoor Attacks Against Hateful Meme Detection via Cross-Modal Triggers
by: Wang, Ruofei, et al.
Published: (2024)
by: Wang, Ruofei, et al.
Published: (2024)
E-SAGE: Explainability-based Defense Against Backdoor Attacks on Graph Neural Networks
by: Yuan, Dingqiang, et al.
Published: (2024)
by: Yuan, Dingqiang, et al.
Published: (2024)
TrojanEdit: Multimodal Backdoor Attack Against Image Editing Model
by: Guo, Ji, et al.
Published: (2024)
by: Guo, Ji, et al.
Published: (2024)
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
by: Ran, Delong, et al.
Published: (2024)
by: Ran, Delong, et al.
Published: (2024)
Revisiting Training-Inference Trigger Intensity in Backdoor Attacks
by: Lin, Chenhao, et al.
Published: (2025)
by: Lin, Chenhao, et al.
Published: (2025)
On the Credibility of Backdoor Attacks Against Object Detectors in the Physical World
by: Doan, Bao Gia, et al.
Published: (2024)
by: Doan, Bao Gia, et al.
Published: (2024)
3S-Attack: Spatial, Spectral and Semantic Invisible Backdoor Attack Against DNN Models
by: Yin, Jianyao, et al.
Published: (2025)
by: Yin, Jianyao, et al.
Published: (2025)
PEFT-as-an-Attack! Jailbreaking Language Models during Federated Parameter-Efficient Fine-Tuning
by: Li, Shenghui, et al.
Published: (2024)
by: Li, Shenghui, et al.
Published: (2024)
Isolate Trigger: Detecting and Eliminating Adaptive Backdoor Attacks
by: Sun, Chengrui, et al.
Published: (2025)
by: Sun, Chengrui, et al.
Published: (2025)
Physical Backdoor Attack Against Deep Learning-Based Modulation Classification
by: Salmi, Younes, et al.
Published: (2026)
by: Salmi, Younes, et al.
Published: (2026)
Instruction Backdoor Attacks Against Customized LLMs
by: Zhang, Rui, et al.
Published: (2024)
by: Zhang, Rui, et al.
Published: (2024)
EmoBack: Backdoor Attacks Against Speaker Identification Using Emotional Prosody
by: Schoof, Coen, et al.
Published: (2024)
by: Schoof, Coen, et al.
Published: (2024)
Fine-tuning is Not Fine: Mitigating Backdoor Attacks in GNNs with Limited Clean Data
by: Zhang, Jiale, et al.
Published: (2025)
by: Zhang, Jiale, et al.
Published: (2025)
Bandwidth-Efficient Two-Server ORAMs with O(1) Client Storage
by: Wang, Wei, et al.
Published: (2025)
by: Wang, Wei, et al.
Published: (2025)
BadMerging: Backdoor Attacks Against Model Merging
by: Zhang, Jinghuai, et al.
Published: (2024)
by: Zhang, Jinghuai, et al.
Published: (2024)
Membership Inference Attack Against Masked Image Modeling
by: Li, Zheng, et al.
Published: (2024)
by: Li, Zheng, et al.
Published: (2024)
Stealthy Backdoor Attack via Confidence-driven Sampling
by: He, Pengfei, et al.
Published: (2023)
by: He, Pengfei, et al.
Published: (2023)
Towards Stealthy and Effective Backdoor Attacks on Lane Detection: A Naturalistic Data Poisoning Approach
by: Liao, Yifan, et al.
Published: (2025)
by: Liao, Yifan, et al.
Published: (2025)
A Chaos Driven Metric for Backdoor Attack Detection
by: Surendrababu, Hema Karnam, et al.
Published: (2025)
by: Surendrababu, Hema Karnam, et al.
Published: (2025)
Can VLMs Detect and Localize Fine-Grained AI-Edited Images?
by: Sun, Zhen, et al.
Published: (2025)
by: Sun, Zhen, et al.
Published: (2025)
Similar Items
-
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
by: Zheng, Jingyi, et al.
Published: (2024) -
Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models
by: Zhang, Zongmin, et al.
Published: (2025) -
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
by: Zhang, Heyi, et al.
Published: (2025) -
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
by: Yi, Sibo, et al.
Published: (2024) -
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
by: Zhao, Shuai, et al.
Published: (2024)