Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Price, Sara, Panickssery, Arjun, Bowman, Sam, Stickland, Asa Cooper |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
von: Sivapiromrat, Sanhanat, et al.
Veröffentlicht: (2025)
von: Sivapiromrat, Sanhanat, et al.
Veröffentlicht: (2025)
Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
Hidden Ads: Behavior Triggered Semantic Backdoors for Advertisement Injection in Vision Language Models
von: Yao, Duanyi, et al.
Veröffentlicht: (2026)
von: Yao, Duanyi, et al.
Veröffentlicht: (2026)
How Vulnerable Are Edge LLMs?
von: Ding, Ao, et al.
Veröffentlicht: (2026)
von: Ding, Ao, et al.
Veröffentlicht: (2026)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
von: Xu, Jiashu, et al.
Veröffentlicht: (2023)
von: Xu, Jiashu, et al.
Veröffentlicht: (2023)
BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers
von: Xue, Jiaqi, et al.
Veröffentlicht: (2024)
von: Xue, Jiaqi, et al.
Veröffentlicht: (2024)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
von: Liu, Qin, et al.
Veröffentlicht: (2023)
von: Liu, Qin, et al.
Veröffentlicht: (2023)
Can LLMs be Fooled? Investigating Vulnerabilities in LLMs
von: Abdali, Sara, et al.
Veröffentlicht: (2024)
von: Abdali, Sara, et al.
Veröffentlicht: (2024)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
von: Xue, Eric, et al.
Veröffentlicht: (2025)
von: Xue, Eric, et al.
Veröffentlicht: (2025)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
von: Tong, Terry, et al.
Veröffentlicht: (2024)
von: Tong, Terry, et al.
Veröffentlicht: (2024)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
von: Rando, Javier, et al.
Veröffentlicht: (2024)
von: Rando, Javier, et al.
Veröffentlicht: (2024)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
von: Betley, Jan, et al.
Veröffentlicht: (2025)
von: Betley, Jan, et al.
Veröffentlicht: (2025)
Composite Backdoor Attacks Against Large Language Models
von: Huang, Hai, et al.
Veröffentlicht: (2023)
von: Huang, Hai, et al.
Veröffentlicht: (2023)
CLIBE: Detecting Dynamic Backdoors in Transformer-based NLP Models
von: Zeng, Rui, et al.
Veröffentlicht: (2024)
von: Zeng, Rui, et al.
Veröffentlicht: (2024)
Data-centric NLP Backdoor Defense from the Lens of Memorization
von: Wang, Zhenting, et al.
Veröffentlicht: (2024)
von: Wang, Zhenting, et al.
Veröffentlicht: (2024)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2024)
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2024)
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
von: Yan, Jun, et al.
Veröffentlicht: (2023)
von: Yan, Jun, et al.
Veröffentlicht: (2023)
Hardware-Triggered Backdoors
von: Möller, Jonas, et al.
Veröffentlicht: (2026)
von: Möller, Jonas, et al.
Veröffentlicht: (2026)
WARDEN: Multi-Directional Backdoor Watermarks for Embedding-as-a-Service Copyright Protection
von: Shetty, Anudeex, et al.
Veröffentlicht: (2024)
von: Shetty, Anudeex, et al.
Veröffentlicht: (2024)
STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio
von: Park, Seong-Gyu, et al.
Veröffentlicht: (2026)
von: Park, Seong-Gyu, et al.
Veröffentlicht: (2026)
Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language Models
von: He, Jiaming, et al.
Veröffentlicht: (2024)
von: He, Jiaming, et al.
Veröffentlicht: (2024)
Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability
von: García-Carrasco, Jorge, et al.
Veröffentlicht: (2024)
von: García-Carrasco, Jorge, et al.
Veröffentlicht: (2024)
Exploring Vulnerabilities and Protections in Large Language Models: A Survey
von: Liu, Frank Weizhen, et al.
Veröffentlicht: (2024)
von: Liu, Frank Weizhen, et al.
Veröffentlicht: (2024)
Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
GCG Attack On A Diffusion LLM
von: Neyroud, Ruben, et al.
Veröffentlicht: (2025)
von: Neyroud, Ruben, et al.
Veröffentlicht: (2025)
WET: Overcoming Paraphrasing Vulnerabilities in Embeddings-as-a-Service with Linear Transformation Watermarks
von: Shetty, Anudeex, et al.
Veröffentlicht: (2024)
von: Shetty, Anudeex, et al.
Veröffentlicht: (2024)
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
von: Doshi, Jai, et al.
Veröffentlicht: (2024)
von: Doshi, Jai, et al.
Veröffentlicht: (2024)
Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluation
von: Guo, Wenkai, et al.
Veröffentlicht: (2025)
von: Guo, Wenkai, et al.
Veröffentlicht: (2025)
The Hidden Cost of Modeling P(X): Vulnerability to Membership Inference Attacks in Generative Text Classifiers
von: Makroo, Owais, et al.
Veröffentlicht: (2025)
von: Makroo, Owais, et al.
Veröffentlicht: (2025)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
Cross-Entropy Attacks to Language Models via Rare Event Simulation
von: Ni, Mingze, et al.
Veröffentlicht: (2025)
von: Ni, Mingze, et al.
Veröffentlicht: (2025)
Coercing LLMs to do and reveal (almost) anything
von: Geiping, Jonas, et al.
Veröffentlicht: (2024)
von: Geiping, Jonas, et al.
Veröffentlicht: (2024)
UCD: Unlearning in LLMs via Contrastive Decoding
von: Suriyakumar, Vinith M., et al.
Veröffentlicht: (2025)
von: Suriyakumar, Vinith M., et al.
Veröffentlicht: (2025)
Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization
von: Dotsinski, Asen, et al.
Veröffentlicht: (2026)
von: Dotsinski, Asen, et al.
Veröffentlicht: (2026)
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
von: Xiong, Alexander, et al.
Veröffentlicht: (2025)
von: Xiong, Alexander, et al.
Veröffentlicht: (2025)
Universal Jailbreak Backdoors from Poisoned Human Feedback
von: Rando, Javier, et al.
Veröffentlicht: (2023)
von: Rando, Javier, et al.
Veröffentlicht: (2023)
Weight space Detection of Backdoors in LoRA Adapters
von: Merenciano, David Puertolas, et al.
Veröffentlicht: (2026)
von: Merenciano, David Puertolas, et al.
Veröffentlicht: (2026)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
von: Wei, Jiali, et al.
Veröffentlicht: (2026)
von: Wei, Jiali, et al.
Veröffentlicht: (2026)
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
von: Wallace, Eric, et al.
Veröffentlicht: (2024)
von: Wallace, Eric, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
von: Sivapiromrat, Sanhanat, et al.
Veröffentlicht: (2025) -
Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities
von: Krishna, Arjun, et al.
Veröffentlicht: (2025) -
Hidden Ads: Behavior Triggered Semantic Backdoors for Advertisement Injection in Vision Language Models
von: Yao, Duanyi, et al.
Veröffentlicht: (2026) -
How Vulnerable Are Edge LLMs?
von: Ding, Ao, et al.
Veröffentlicht: (2026) -
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
von: Xu, Jiashu, et al.
Veröffentlicht: (2023)