Data-centric NLP Backdoor Defense from the Lens of Memorization
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Zhenting, Wang, Zhizhi, Jin, Mingyu, Du, Mengnan, Zhai, Juan, Ma, Shiqing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CLIBE: Detecting Dynamic Backdoors in Transformer-based NLP Models
di: Zeng, Rui, et al.
Pubblicazione: (2024)
di: Zeng, Rui, et al.
Pubblicazione: (2024)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
di: Xu, Jiashu, et al.
Pubblicazione: (2023)
di: Xu, Jiashu, et al.
Pubblicazione: (2023)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
di: Liu, Qin, et al.
Pubblicazione: (2023)
di: Liu, Qin, et al.
Pubblicazione: (2023)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
di: Xue, Eric, et al.
Pubblicazione: (2025)
di: Xue, Eric, et al.
Pubblicazione: (2025)
Leaner Training, Lower Leakage: Revisiting Memorization in LLM Fine-Tuning with LoRA
di: Wang, Fei, et al.
Pubblicazione: (2025)
di: Wang, Fei, et al.
Pubblicazione: (2025)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
di: Zeng, Yifan, et al.
Pubblicazione: (2024)
di: Zeng, Yifan, et al.
Pubblicazione: (2024)
EmojiPrompt: Generative Prompt Obfuscation for Privacy-Preserving Communication with Cloud-based LLMs
di: Lin, Sam, et al.
Pubblicazione: (2024)
di: Lin, Sam, et al.
Pubblicazione: (2024)
DIAGNOSIS: Detecting Unauthorized Data Usages in Text-to-image Diffusion Models
di: Wang, Zhenting, et al.
Pubblicazione: (2023)
di: Wang, Zhenting, et al.
Pubblicazione: (2023)
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
di: Yan, Jun, et al.
Pubblicazione: (2023)
di: Yan, Jun, et al.
Pubblicazione: (2023)
Localizing Paragraph Memorization in Language Models
di: Stoehr, Niklas, et al.
Pubblicazione: (2024)
di: Stoehr, Niklas, et al.
Pubblicazione: (2024)
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluation
di: Guo, Wenkai, et al.
Pubblicazione: (2025)
di: Guo, Wenkai, et al.
Pubblicazione: (2025)
Randomized Masked Finetuning: An Efficient Way to Mitigate Memorization of PIIs in LLMs
di: Joshi, Kunj, et al.
Pubblicazione: (2025)
di: Joshi, Kunj, et al.
Pubblicazione: (2025)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
di: Qi, Zhenting, et al.
Pubblicazione: (2024)
di: Qi, Zhenting, et al.
Pubblicazione: (2024)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
di: Sivapiromrat, Sanhanat, et al.
Pubblicazione: (2025)
di: Sivapiromrat, Sanhanat, et al.
Pubblicazione: (2025)
Composite Backdoor Attacks Against Large Language Models
di: Huang, Hai, et al.
Pubblicazione: (2023)
di: Huang, Hai, et al.
Pubblicazione: (2023)
Semantic Stealth: Adversarial Text Attacks on NLP Using Several Methods
di: Dey, Roopkatha, et al.
Pubblicazione: (2024)
di: Dey, Roopkatha, et al.
Pubblicazione: (2024)
On the Privacy Effect of Data Enhancement via the Lens of Memorization
di: Li, Xiao, et al.
Pubblicazione: (2022)
di: Li, Xiao, et al.
Pubblicazione: (2022)
TwinGate: Stateful Defense against Decompositional Jailbreaks in Untraceable Traffic via Asymmetric Contrastive Learning
di: Sun, Bowen, et al.
Pubblicazione: (2026)
di: Sun, Bowen, et al.
Pubblicazione: (2026)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
di: Chaudhari, Harsh, et al.
Pubblicazione: (2024)
di: Chaudhari, Harsh, et al.
Pubblicazione: (2024)
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
di: Price, Sara, et al.
Pubblicazione: (2024)
di: Price, Sara, et al.
Pubblicazione: (2024)
From Theory to Practice: Evaluating Data Poisoning Attacks and Defenses in In-Context Learning on Social Media Health Discourse
di: Jhuma, Rabeya Amin, et al.
Pubblicazione: (2025)
di: Jhuma, Rabeya Amin, et al.
Pubblicazione: (2025)
Robustness Inspired Graph Backdoor Defense
di: Zhang, Zhiwei, et al.
Pubblicazione: (2024)
di: Zhang, Zhiwei, et al.
Pubblicazione: (2024)
Testing the Limits of Jailbreaking Defenses with the Purple Problem
di: Kim, Taeyoun, et al.
Pubblicazione: (2024)
di: Kim, Taeyoun, et al.
Pubblicazione: (2024)
Position: Privacy Is Not Just Memorization!
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2025)
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2025)
WARDEN: Multi-Directional Backdoor Watermarks for Embedding-as-a-Service Copyright Protection
di: Shetty, Anudeex, et al.
Pubblicazione: (2024)
di: Shetty, Anudeex, et al.
Pubblicazione: (2024)
Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses
di: Ahmed, Mohamed, et al.
Pubblicazione: (2025)
di: Ahmed, Mohamed, et al.
Pubblicazione: (2025)
Unlocking Memorization in Large Language Models with Dynamic Soft Prompting
di: Wang, Zhepeng, et al.
Pubblicazione: (2024)
di: Wang, Zhepeng, et al.
Pubblicazione: (2024)
Hidden Ads: Behavior Triggered Semantic Backdoors for Advertisement Injection in Vision Language Models
di: Yao, Duanyi, et al.
Pubblicazione: (2026)
di: Yao, Duanyi, et al.
Pubblicazione: (2026)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
di: Brown, Hannah, et al.
Pubblicazione: (2024)
di: Brown, Hannah, et al.
Pubblicazione: (2024)
Gradient Shaping: Enhancing Backdoor Attack Against Reverse Engineering
di: Zhu, Rui, et al.
Pubblicazione: (2023)
di: Zhu, Rui, et al.
Pubblicazione: (2023)
STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio
di: Park, Seong-Gyu, et al.
Pubblicazione: (2026)
di: Park, Seong-Gyu, et al.
Pubblicazione: (2026)
Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models
di: Dong, Yihong, et al.
Pubblicazione: (2024)
di: Dong, Yihong, et al.
Pubblicazione: (2024)
Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language Models
di: He, Jiaming, et al.
Pubblicazione: (2024)
di: He, Jiaming, et al.
Pubblicazione: (2024)
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
di: Kim, Heegyu, et al.
Pubblicazione: (2024)
di: Kim, Heegyu, et al.
Pubblicazione: (2024)
Subspace Defense: Discarding Adversarial Perturbations by Learning a Subspace for Clean Signals
di: Zheng, Rui, et al.
Pubblicazione: (2024)
di: Zheng, Rui, et al.
Pubblicazione: (2024)
On the Detectability of ChatGPT Content: Benchmarking, Methodology, and Evaluation through the Lens of Academic Writing
di: Liu, Zeyan, et al.
Pubblicazione: (2023)
di: Liu, Zeyan, et al.
Pubblicazione: (2023)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
di: Wang, Yifei, et al.
Pubblicazione: (2024)
di: Wang, Yifei, et al.
Pubblicazione: (2024)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
di: Li, Nathaniel, et al.
Pubblicazione: (2024)
di: Li, Nathaniel, et al.
Pubblicazione: (2024)
Seal Your Backdoor with Variational Defense
di: Sabolić, Ivan, et al.
Pubblicazione: (2025)
di: Sabolić, Ivan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CLIBE: Detecting Dynamic Backdoors in Transformer-based NLP Models
di: Zeng, Rui, et al.
Pubblicazione: (2024) -
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
di: Xu, Jiashu, et al.
Pubblicazione: (2023) -
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
di: Liu, Qin, et al.
Pubblicazione: (2023) -
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
di: Xue, Eric, et al.
Pubblicazione: (2025) -
Leaner Training, Lower Leakage: Revisiting Memorization in LLM Fine-Tuning with LoRA
di: Wang, Fei, et al.
Pubblicazione: (2025)