Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Betley, Jan, Cocola, Jorio, Feng, Dylan, Chua, James, Arditi, Andy, Sztyber-Betley, Anna, Evans, Owain |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tell me about yourself: LLMs are aware of their learned behaviors
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
by: Dubiński, Jan, et al.
Published: (2026)
by: Dubiński, Jan, et al.
Published: (2026)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
by: Chua, James, et al.
Published: (2025)
by: Chua, James, et al.
Published: (2025)
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
by: Chua, James, et al.
Published: (2026)
by: Chua, James, et al.
Published: (2026)
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
by: Cloud, Alex, et al.
Published: (2025)
by: Cloud, Alex, et al.
Published: (2025)
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
by: Taylor, Mia, et al.
Published: (2025)
by: Taylor, Mia, et al.
Published: (2025)
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
by: Wen, Rui, et al.
Published: (2026)
by: Wen, Rui, et al.
Published: (2026)
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
by: Zhao, Gejian, et al.
Published: (2025)
by: Zhao, Gejian, et al.
Published: (2025)
TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning
by: He, Xuanli, et al.
Published: (2024)
by: He, Xuanli, et al.
Published: (2024)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Privacy Backdoors: Stealing Data with Corrupted Pretrained Models
by: Feng, Shanglun, et al.
Published: (2024)
by: Feng, Shanglun, et al.
Published: (2024)
Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models
by: Li, Haoran, et al.
Published: (2024)
by: Li, Haoran, et al.
Published: (2024)
Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors
by: Peng, Yuefeng, et al.
Published: (2024)
by: Peng, Yuefeng, et al.
Published: (2024)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
by: Ji, Wence, et al.
Published: (2025)
by: Ji, Wence, et al.
Published: (2025)
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
by: Price, Sara, et al.
Published: (2024)
by: Price, Sara, et al.
Published: (2024)
TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
by: Cheng, Pengzhou, et al.
Published: (2024)
by: Cheng, Pengzhou, et al.
Published: (2024)
AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption
by: Wang, Yanting, et al.
Published: (2025)
by: Wang, Yanting, et al.
Published: (2025)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
by: Wei, Jiali, et al.
Published: (2026)
by: Wei, Jiali, et al.
Published: (2026)
Rethinking Backdoor Detection Evaluation for Language Models
by: Yan, Jun, et al.
Published: (2024)
by: Yan, Jun, et al.
Published: (2024)
Task-Agnostic Detector for Insertion-Based Backdoor Attacks
by: Lyu, Weimin, et al.
Published: (2024)
by: Lyu, Weimin, et al.
Published: (2024)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
by: Chaudhari, Harsh, et al.
Published: (2024)
by: Chaudhari, Harsh, et al.
Published: (2024)
Lightweight Yet Secure: Secure Scripting Language Generation via Lightweight LLMs
by: Zhang, Keyang, et al.
Published: (2026)
by: Zhang, Keyang, et al.
Published: (2026)
P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs
by: Zhao, Shuai, et al.
Published: (2025)
by: Zhao, Shuai, et al.
Published: (2025)
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
BadActs: A Universal Backdoor Defense in the Activation Space
by: Yi, Biao, et al.
Published: (2024)
by: Yi, Biao, et al.
Published: (2024)
Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models
by: Fortier, Alexandrine, et al.
Published: (2025)
by: Fortier, Alexandrine, et al.
Published: (2025)
Cut the Deadwood Out: Backdoor Purification via Guided Module Substitution
by: Tong, Yao, et al.
Published: (2024)
by: Tong, Yao, et al.
Published: (2024)
LLMs for Domain Generation Algorithm Detection
by: La O, Reynier Leyva, et al.
Published: (2024)
by: La O, Reynier Leyva, et al.
Published: (2024)
Randomized Masked Finetuning: An Efficient Way to Mitigate Memorization of PIIs in LLMs
by: Joshi, Kunj, et al.
Published: (2025)
by: Joshi, Kunj, et al.
Published: (2025)
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
by: Treutlein, Johannes, et al.
Published: (2024)
by: Treutlein, Johannes, et al.
Published: (2024)
ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Compiling Activation Steering into Weights via Null-Space Constraints for Stealthy Backdoors
by: Yin, Rui, et al.
Published: (2026)
by: Yin, Rui, et al.
Published: (2026)
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
by: Li, Ziqiang, et al.
Published: (2024)
by: Li, Ziqiang, et al.
Published: (2024)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
by: Wang, Jiongxiao, et al.
Published: (2024)
by: Wang, Jiongxiao, et al.
Published: (2024)
MBTSAD: Mitigating Backdoors in Language Models Based on Token Splitting and Attention Distillation
by: Ding, Yidong, et al.
Published: (2025)
by: Ding, Yidong, et al.
Published: (2025)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
by: Xue, Eric, et al.
Published: (2025)
by: Xue, Eric, et al.
Published: (2025)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
by: Jiang, Peihai, et al.
Published: (2025)
by: Jiang, Peihai, et al.
Published: (2025)
SEEP: Training Dynamics Grounds Latent Representation Search for Mitigating Backdoor Poisoning Attacks
by: He, Xuanli, et al.
Published: (2024)
by: He, Xuanli, et al.
Published: (2024)
Similar Items
-
Tell me about yourself: LLMs are aware of their learned behaviors
by: Betley, Jan, et al.
Published: (2025) -
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
by: Dubiński, Jan, et al.
Published: (2026) -
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
by: Betley, Jan, et al.
Published: (2025) -
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
by: Chua, James, et al.
Published: (2025) -
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
by: Chua, James, et al.
Published: (2026)