A Study of Backdoors in Instruction Fine-tuned Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Raghuram, Jayaram, Kesidis, George, Miller, David J. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
von: Xu, Jiashu, et al.
Veröffentlicht: (2023)
von: Xu, Jiashu, et al.
Veröffentlicht: (2023)
Obliviate: Neutralizing Task-agnostic Backdoors within the Parameter-efficient Fine-tuning Paradigm
von: Kim, Jaehan, et al.
Veröffentlicht: (2024)
von: Kim, Jaehan, et al.
Veröffentlicht: (2024)
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
von: Huang, Tiansheng, et al.
Veröffentlicht: (2025)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2025)
SentinelLMs: Encrypted Input Adaptation and Fine-tuning of Language Models for Private and Secure Inference
von: Mishra, Abhijit, et al.
Veröffentlicht: (2023)
von: Mishra, Abhijit, et al.
Veröffentlicht: (2023)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
von: Tong, Terry, et al.
Veröffentlicht: (2024)
von: Tong, Terry, et al.
Veröffentlicht: (2024)
Instructional Fingerprinting of Large Language Models
von: Xu, Jiashu, et al.
Veröffentlicht: (2024)
von: Xu, Jiashu, et al.
Veröffentlicht: (2024)
ObfuscaTune: Obfuscated Offsite Fine-tuning and Inference of Proprietary LLMs on Private Datasets
von: Frikha, Ahmed, et al.
Veröffentlicht: (2024)
von: Frikha, Ahmed, et al.
Veröffentlicht: (2024)
Weight space Detection of Backdoors in LoRA Adapters
von: Merenciano, David Puertolas, et al.
Veröffentlicht: (2026)
von: Merenciano, David Puertolas, et al.
Veröffentlicht: (2026)
CEPA: Consensus Embedded Perturbation for Agnostic Detection and Inversion of Backdoors
von: Yang, Guangmingmei, et al.
Veröffentlicht: (2024)
von: Yang, Guangmingmei, et al.
Veröffentlicht: (2024)
Fine-Tuning Language Models with Differential Privacy through Adaptive Noise Allocation
von: Li, Xianzhi, et al.
Veröffentlicht: (2024)
von: Li, Xianzhi, et al.
Veröffentlicht: (2024)
The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2023)
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2023)
Universal Jailbreak Backdoors from Poisoned Human Feedback
von: Rando, Javier, et al.
Veröffentlicht: (2023)
von: Rando, Javier, et al.
Veröffentlicht: (2023)
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
von: Liu, Qin, et al.
Veröffentlicht: (2024)
von: Liu, Qin, et al.
Veröffentlicht: (2024)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
von: Rando, Javier, et al.
Veröffentlicht: (2024)
von: Rando, Javier, et al.
Veröffentlicht: (2024)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
von: Liu, Qin, et al.
Veröffentlicht: (2023)
von: Liu, Qin, et al.
Veröffentlicht: (2023)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
von: Betley, Jan, et al.
Veröffentlicht: (2025)
von: Betley, Jan, et al.
Veröffentlicht: (2025)
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
von: Zhang, Zhixin, et al.
Veröffentlicht: (2025)
von: Zhang, Zhixin, et al.
Veröffentlicht: (2025)
MEUV: Achieving Fine-Grained Capability Activation in Large Language Models via Mutually Exclusive Unlock Vectors
von: Tong, Xin, et al.
Veröffentlicht: (2025)
von: Tong, Xin, et al.
Veröffentlicht: (2025)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
von: Wu, Tong, et al.
Veröffentlicht: (2024)
von: Wu, Tong, et al.
Veröffentlicht: (2024)
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
von: Yan, Jun, et al.
Veröffentlicht: (2023)
von: Yan, Jun, et al.
Veröffentlicht: (2023)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
Memories Retrieved from Many Paths: A Multi-Prefix Framework for Robust Detection of Training Data Leakage in Large Language Models
von: Dang, Trung Cuong, et al.
Veröffentlicht: (2025)
von: Dang, Trung Cuong, et al.
Veröffentlicht: (2025)
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning
von: Wang, Shanmin, et al.
Veröffentlicht: (2025)
von: Wang, Shanmin, et al.
Veröffentlicht: (2025)
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
von: Guo, Chuan, et al.
Veröffentlicht: (2026)
von: Guo, Chuan, et al.
Veröffentlicht: (2026)
Adaptive Instruction Composition for Automated LLM Red-Teaming
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
von: Jiang, Peihai, et al.
Veröffentlicht: (2025)
von: Jiang, Peihai, et al.
Veröffentlicht: (2025)
Private Fine-tuning of Large Language Models with Zeroth-order Optimization
von: Tang, Xinyu, et al.
Veröffentlicht: (2024)
von: Tang, Xinyu, et al.
Veröffentlicht: (2024)
Window-based Membership Inference Attacks Against Fine-tuned Large Language Models
von: Chen, Yuetian, et al.
Veröffentlicht: (2026)
von: Chen, Yuetian, et al.
Veröffentlicht: (2026)
Watermarking Makes Language Models Radioactive
von: Sander, Tom, et al.
Veröffentlicht: (2024)
von: Sander, Tom, et al.
Veröffentlicht: (2024)
Lifelong Safety Alignment for Language Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Personal Information Parroting in Language Models
von: Subramani, Nishant, et al.
Veröffentlicht: (2026)
von: Subramani, Nishant, et al.
Veröffentlicht: (2026)
PostMark: A Robust Blackbox Watermark for Large Language Models
von: Chang, Yapei, et al.
Veröffentlicht: (2024)
von: Chang, Yapei, et al.
Veröffentlicht: (2024)
Adversarial Text Purification: A Large Language Model Approach for Defense
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
GaussMark: A Practical Approach for Structural Watermarking of Language Models
von: Block, Adam, et al.
Veröffentlicht: (2025)
von: Block, Adam, et al.
Veröffentlicht: (2025)
Preserving Privacy in Large Language Models: A Survey on Current Threats and Solutions
von: Miranda, Michele, et al.
Veröffentlicht: (2024)
von: Miranda, Michele, et al.
Veröffentlicht: (2024)
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
von: Li, Lijun, et al.
Veröffentlicht: (2024)
von: Li, Lijun, et al.
Veröffentlicht: (2024)
RECAP: A Resource-Efficient Method for Adversarial Prompting in Large Language Models
von: Chugh, Rishit
Veröffentlicht: (2026)
von: Chugh, Rishit
Veröffentlicht: (2026)
Ähnliche Einträge
-
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
von: Xu, Jiashu, et al.
Veröffentlicht: (2023) -
Obliviate: Neutralizing Task-agnostic Backdoors within the Parameter-efficient Fine-tuning Paradigm
von: Kim, Jaehan, et al.
Veröffentlicht: (2024) -
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
von: Huang, Tiansheng, et al.
Veröffentlicht: (2025) -
SentinelLMs: Encrypted Input Adaptation and Fine-tuning of Language Models for Private and Secure Inference
von: Mishra, Abhijit, et al.
Veröffentlicht: (2023) -
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
von: Tong, Terry, et al.
Veröffentlicht: (2024)