Detecting Instruction Fine-tuning Attacks using Influence Function
Fuente:
arXiv
Salvato in:
| Autore principale: | Li, Jiawei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Fine-tuning is Not Fine: Mitigating Backdoor Attacks in GNNs with Limited Clean Data
di: Zhang, Jiale, et al.
Pubblicazione: (2025)
di: Zhang, Jiale, et al.
Pubblicazione: (2025)
SoK: Reducing the Vulnerability of Fine-tuned Language Models to Membership Inference Attacks
di: Amit, Guy, et al.
Pubblicazione: (2024)
di: Amit, Guy, et al.
Pubblicazione: (2024)
Selective Pre-training for Private Fine-tuning
di: Yu, Da, et al.
Pubblicazione: (2023)
di: Yu, Da, et al.
Pubblicazione: (2023)
Can Differentially Private Fine-tuning LLMs Protect Against Privacy Attacks?
di: Du, Hao, et al.
Pubblicazione: (2025)
di: Du, Hao, et al.
Pubblicazione: (2025)
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
di: Huang, Tiansheng, et al.
Pubblicazione: (2024)
di: Huang, Tiansheng, et al.
Pubblicazione: (2024)
Instruction Backdoor Attacks Against Customized LLMs
di: Zhang, Rui, et al.
Pubblicazione: (2024)
di: Zhang, Rui, et al.
Pubblicazione: (2024)
Shake to Leak: Fine-tuning Diffusion Models Can Amplify the Generative Privacy Risk
di: Li, Zhangheng, et al.
Pubblicazione: (2024)
di: Li, Zhangheng, et al.
Pubblicazione: (2024)
AttackQA: Development and Adoption of a Dataset for Assisting Cybersecurity Operations using Fine-tuned and Open-Source LLMs
di: Krishna, Varun Badrinath
Pubblicazione: (2024)
di: Krishna, Varun Badrinath
Pubblicazione: (2024)
I can't see it but I can Fine-tune it: On Encrypted Fine-tuning of Transformers using Fully Homomorphic Encryption
di: Panzade, Prajwal, et al.
Pubblicazione: (2024)
di: Panzade, Prajwal, et al.
Pubblicazione: (2024)
Attack Smarter: Attention-Driven Fine-Grained Webpage Fingerprinting Attacks
di: Yuan, Yali, et al.
Pubblicazione: (2025)
di: Yuan, Yali, et al.
Pubblicazione: (2025)
Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
di: Fu, Wenjie, et al.
Pubblicazione: (2023)
di: Fu, Wenjie, et al.
Pubblicazione: (2023)
Navigating the Designs of Privacy-Preserving Fine-tuning for Large Language Models
di: Shi, Haonan, et al.
Pubblicazione: (2025)
di: Shi, Haonan, et al.
Pubblicazione: (2025)
Fine-tuning of Large Language Models for Domain-Specific Cybersecurity Knowledge
di: Huang, Yuan
Pubblicazione: (2025)
di: Huang, Yuan
Pubblicazione: (2025)
FRIDA: Free-Rider Detection using Privacy Attacks
di: Recasens, Pol G., et al.
Pubblicazione: (2024)
di: Recasens, Pol G., et al.
Pubblicazione: (2024)
A Study of Backdoors in Instruction Fine-tuned Language Models
di: Raghuram, Jayaram, et al.
Pubblicazione: (2024)
di: Raghuram, Jayaram, et al.
Pubblicazione: (2024)
UniASM: Binary Code Similarity Detection without Fine-tuning
di: Gu, Yeming, et al.
Pubblicazione: (2022)
di: Gu, Yeming, et al.
Pubblicazione: (2022)
Private LoRA Fine-tuning of Open-Source LLMs with Homomorphic Encryption
di: Frery, Jordan, et al.
Pubblicazione: (2025)
di: Frery, Jordan, et al.
Pubblicazione: (2025)
Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data
di: Akkus, Atilla, et al.
Pubblicazione: (2024)
di: Akkus, Atilla, et al.
Pubblicazione: (2024)
GasTrace: Detecting Sandwich Attack Malicious Accounts in Ethereum
di: Liu, Zekai, et al.
Pubblicazione: (2024)
di: Liu, Zekai, et al.
Pubblicazione: (2024)
Using Anomaly Detection to Detect Poisoning Attacks in Federated Learning Applications
di: Raza, Ali, et al.
Pubblicazione: (2022)
di: Raza, Ali, et al.
Pubblicazione: (2022)
Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks
di: Kuo, Kevin, et al.
Pubblicazione: (2026)
di: Kuo, Kevin, et al.
Pubblicazione: (2026)
Double-I Watermark: Protecting Model Copyright for LLM Fine-tuning
di: Li, Shen, et al.
Pubblicazione: (2024)
di: Li, Shen, et al.
Pubblicazione: (2024)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
di: Hsiung, Lei, et al.
Pubblicazione: (2025)
di: Hsiung, Lei, et al.
Pubblicazione: (2025)
Attack and Defense of Deep Learning Models in the Field of Web Attack Detection
di: Shi, Lijia, et al.
Pubblicazione: (2024)
di: Shi, Lijia, et al.
Pubblicazione: (2024)
Rethinking PGD Attack: Is Sign Function Necessary?
di: Yang, Junjie, et al.
Pubblicazione: (2023)
di: Yang, Junjie, et al.
Pubblicazione: (2023)
LLMStinger: Jailbreaking LLMs using RL fine-tuned LLMs
di: Jha, Piyush, et al.
Pubblicazione: (2024)
di: Jha, Piyush, et al.
Pubblicazione: (2024)
ADVENT: Attack/Anomaly Detection in VANETs
di: Baharlouei, Hamideh, et al.
Pubblicazione: (2024)
di: Baharlouei, Hamideh, et al.
Pubblicazione: (2024)
Fine Grained Insider Risk Detection
di: Huber, Birkett, et al.
Pubblicazione: (2024)
di: Huber, Birkett, et al.
Pubblicazione: (2024)
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
di: Huang, Tiansheng, et al.
Pubblicazione: (2025)
di: Huang, Tiansheng, et al.
Pubblicazione: (2025)
Delta-Influence: Unlearning Poisons via Influence Functions
di: Li, Wenjie, et al.
Pubblicazione: (2024)
di: Li, Wenjie, et al.
Pubblicazione: (2024)
Privately Learning from Graphs with Applications in Fine-tuning Large Language Models
di: Yin, Haoteng, et al.
Pubblicazione: (2024)
di: Yin, Haoteng, et al.
Pubblicazione: (2024)
IF-GUIDE: Influence Function-Guided Detoxification of LLMs
di: Coalson, Zachary, et al.
Pubblicazione: (2025)
di: Coalson, Zachary, et al.
Pubblicazione: (2025)
Heterogeneous Graph Backdoor Attack
di: Chen, Jiawei, et al.
Pubblicazione: (2025)
di: Chen, Jiawei, et al.
Pubblicazione: (2025)
Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
di: Liu, Guozhi, et al.
Pubblicazione: (2025)
di: Liu, Guozhi, et al.
Pubblicazione: (2025)
Automated Membership Inference Attacks: Discovering MIA Signal Computations using LLM Agents
di: Tran, Toan, et al.
Pubblicazione: (2026)
di: Tran, Toan, et al.
Pubblicazione: (2026)
Analysis of Zero Day Attack Detection Using MLP and XAI
di: Dahal, Ashim, et al.
Pubblicazione: (2025)
di: Dahal, Ashim, et al.
Pubblicazione: (2025)
Detecting Backdoor Attacks via Similarity in Semantic Communication Systems
di: Wei, Ziyang, et al.
Pubblicazione: (2025)
di: Wei, Ziyang, et al.
Pubblicazione: (2025)
DMGNN: Detecting and Mitigating Backdoor Attacks in Graph Neural Networks
di: Sui, Hao, et al.
Pubblicazione: (2024)
di: Sui, Hao, et al.
Pubblicazione: (2024)
SENet: Visual Detection of Online Social Engineering Attack Campaigns
di: Ozen, Irfan, et al.
Pubblicazione: (2024)
di: Ozen, Irfan, et al.
Pubblicazione: (2024)
Adaptive Anomaly Detection for Identifying Attacks in Cyber-Physical Systems: A Systematic Literature Review
di: Moriano, Pablo, et al.
Pubblicazione: (2024)
di: Moriano, Pablo, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Fine-tuning is Not Fine: Mitigating Backdoor Attacks in GNNs with Limited Clean Data
di: Zhang, Jiale, et al.
Pubblicazione: (2025) -
SoK: Reducing the Vulnerability of Fine-tuned Language Models to Membership Inference Attacks
di: Amit, Guy, et al.
Pubblicazione: (2024) -
Selective Pre-training for Private Fine-tuning
di: Yu, Da, et al.
Pubblicazione: (2023) -
Can Differentially Private Fine-tuning LLMs Protect Against Privacy Attacks?
di: Du, Hao, et al.
Pubblicazione: (2025) -
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
di: Huang, Tiansheng, et al.
Pubblicazione: (2024)