Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Poppi, Samuele, Yong, Zheng-Xin, He, Yifei, Chern, Bobbie, Zhao, Han, Yang, Aobo, Chi, Jianfeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs
by: Cui, Jing, et al.
Published: (2025)
by: Cui, Jing, et al.
Published: (2025)
PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-Tuning
by: Sun, Zhen, et al.
Published: (2024)
by: Sun, Zhen, et al.
Published: (2024)
VortexPIA: Indirect Prompt Injection Attack against LLMs for Efficient Extraction of User Privacy
by: Cui, Yu, et al.
Published: (2025)
by: Cui, Yu, et al.
Published: (2025)
Do Fine-Tuned LLMs Understand Vulnerabilities? An Investigation into the Semantic Trap
by: Huang, Feiyang, et al.
Published: (2026)
by: Huang, Feiyang, et al.
Published: (2026)
Your Semantic-Independent Watermark is Fragile: A Semantic Perturbation Attack against EaaS Watermark
by: Fei, Zekun, et al.
Published: (2024)
by: Fei, Zekun, et al.
Published: (2024)
ShadowCode: Towards (Automatic) External Prompt Injection Attack against Code LLMs
by: Yang, Yuchen, et al.
Published: (2024)
by: Yang, Yuchen, et al.
Published: (2024)
External Data Extraction Attacks against Retrieval-Augmented Large Language Models
by: He, Yu, et al.
Published: (2025)
by: He, Yu, et al.
Published: (2025)
Scam Shield: Multi-Model Voting and Fine-Tuned LLMs Against Adversarial Attacks
by: Chang, Chen-Wei, et al.
Published: (2025)
by: Chang, Chen-Wei, et al.
Published: (2025)
AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding
by: Ma, Haokai, et al.
Published: (2025)
by: Ma, Haokai, et al.
Published: (2025)
Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment
by: Kim, Jaehan, et al.
Published: (2025)
by: Kim, Jaehan, et al.
Published: (2025)
Black-box Membership Inference Attacks against Fine-tuned Diffusion Models
by: Pang, Yan, et al.
Published: (2023)
by: Pang, Yan, et al.
Published: (2023)
Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
by: Zhao, Shiqian, et al.
Published: (2025)
by: Zhao, Shiqian, et al.
Published: (2025)
Towards a Systematic Taxonomy of Attacks against Space Infrastructures
by: Remy, Jose Luis Castanon, et al.
Published: (2025)
by: Remy, Jose Luis Castanon, et al.
Published: (2025)
Towards Unveiling Vulnerabilities of Large Reasoning Models in Machine Unlearning
by: Chen, Aobo, et al.
Published: (2026)
by: Chen, Aobo, et al.
Published: (2026)
VIMU: Effective Physics-based Realtime Detection and Recovery against Stealthy Attacks on UAVs
by: Wang, Yunbo, et al.
Published: (2025)
by: Wang, Yunbo, et al.
Published: (2025)
Adversarial Attack Based Countermeasures against Deep Learning Side-Channel Attacks
by: Gu, Ruizhe, et al.
Published: (2020)
by: Gu, Ruizhe, et al.
Published: (2020)
Combinational Backdoor Attack against Customized Text-to-Image Models
by: Jiang, Wenbo, et al.
Published: (2024)
by: Jiang, Wenbo, et al.
Published: (2024)
Robust Safety Monitoring of Language Models via Activation Watermarking
by: Aremu, Toluwani, et al.
Published: (2026)
by: Aremu, Toluwani, et al.
Published: (2026)
Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
by: Xie, Zhixin, et al.
Published: (2025)
by: Xie, Zhixin, et al.
Published: (2025)
Hijacking Attacks against Neural Networks by Analyzing Training Data
by: Ge, Yunjie, et al.
Published: (2024)
by: Ge, Yunjie, et al.
Published: (2024)
BadMerging: Backdoor Attacks Against Model Merging
by: Zhang, Jinghuai, et al.
Published: (2024)
by: Zhang, Jinghuai, et al.
Published: (2024)
Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
by: Fu, Haowei, et al.
Published: (2025)
by: Fu, Haowei, et al.
Published: (2025)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
by: Xiong, Chen, et al.
Published: (2024)
by: Xiong, Chen, et al.
Published: (2024)
RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMs
by: Chen, Xuan, et al.
Published: (2024)
by: Chen, Xuan, et al.
Published: (2024)
Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
Demonstration Attack against In-Context Learning for Code Intelligence
by: Ge, Yifei, et al.
Published: (2024)
by: Ge, Yifei, et al.
Published: (2024)
Reference Recommendation based Membership Inference Attack against Hybrid-based Recommender Systems
by: Chi, Xiaoxiao, et al.
Published: (2025)
by: Chi, Xiaoxiao, et al.
Published: (2025)
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
by: Roh, Jaechul, et al.
Published: (2026)
by: Roh, Jaechul, et al.
Published: (2026)
PINA: Prompt Injection Attack against Navigation Agents
by: Liu, Jiani, et al.
Published: (2026)
by: Liu, Jiani, et al.
Published: (2026)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning
by: He, Xuanli, et al.
Published: (2024)
by: He, Xuanli, et al.
Published: (2024)
Fun-tuning: Characterizing the Vulnerability of Proprietary LLMs to Optimization-based Prompt Injection Attacks via the Fine-Tuning Interface
by: Labunets, Andrey, et al.
Published: (2025)
by: Labunets, Andrey, et al.
Published: (2025)
Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes
by: Kumar, Divyanshu, et al.
Published: (2024)
by: Kumar, Divyanshu, et al.
Published: (2024)
When FinTech Meets Privacy: Securing Financial LLMs with Differential Private Fine-Tuning
by: Zhu, Sichen, et al.
Published: (2025)
by: Zhu, Sichen, et al.
Published: (2025)
Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors
by: Yueh-Han, Chen, et al.
Published: (2025)
by: Yueh-Han, Chen, et al.
Published: (2025)
DF-LoGiT: Data-Free Logic-Gated Backdoor Attacks in Vision Transformers
by: Shen, Xiaozuo, et al.
Published: (2026)
by: Shen, Xiaozuo, et al.
Published: (2026)
StrTune: Data Dependence-based Code Slicing for Binary Similarity Detection with Fine-tuned Representation
by: He, Kaiyan, et al.
Published: (2024)
by: He, Kaiyan, et al.
Published: (2024)
Enhancing Privacy of Spatiotemporal Federated Learning against Gradient Inversion Attacks
by: Zheng, Lele, et al.
Published: (2024)
by: Zheng, Lele, et al.
Published: (2024)
Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks
by: Kuo, Kevin, et al.
Published: (2026)
by: Kuo, Kevin, et al.
Published: (2026)
A Stealthy Wrongdoer: Feature-Oriented Reconstruction Attack against Split Learning
by: Xu, Xiaoyang, et al.
Published: (2024)
by: Xu, Xiaoyang, et al.
Published: (2024)
Similar Items
-
Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs
by: Cui, Jing, et al.
Published: (2025) -
PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-Tuning
by: Sun, Zhen, et al.
Published: (2024) -
VortexPIA: Indirect Prompt Injection Attack against LLMs for Efficient Extraction of User Privacy
by: Cui, Yu, et al.
Published: (2025) -
Do Fine-Tuned LLMs Understand Vulnerabilities? An Investigation into the Semantic Trap
by: Huang, Feiyang, et al.
Published: (2026) -
Your Semantic-Independent Watermark is Fragile: A Semantic Perturbation Attack against EaaS Watermark
by: Fei, Zekun, et al.
Published: (2024)