You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation
Fuente:
arXiv
Saved in:
| Main Authors: | Arazzi, Marco, Kembu, Vignesh Kumar, Nocera, Antonino, Picek, Stjepan, Sakthidharan, Saraga |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SecureBreak -- A dataset towards safe and secure models
by: Arazzi, Marco, et al.
Published: (2026)
by: Arazzi, Marco, et al.
Published: (2026)
Security in LLM-as-a-Judge: A Comprehensive SoK
by: Masoud, Aiman Al, et al.
Published: (2026)
by: Masoud, Aiman Al, et al.
Published: (2026)
XBreaking: Understanding how LLMs security alignment can be broken
by: Arazzi, Marco, et al.
Published: (2025)
by: Arazzi, Marco, et al.
Published: (2025)
Let's Focus: Focused Backdoor Attack against Federated Transfer Learning
by: Arazzi, Marco, et al.
Published: (2024)
by: Arazzi, Marco, et al.
Published: (2024)
LoRA as Oracle
by: Arazzi, Marco, et al.
Published: (2026)
by: Arazzi, Marco, et al.
Published: (2026)
Label Inference Attacks against Node-level Vertical Federated GNNs
by: Arazzi, Marco, et al.
Published: (2023)
by: Arazzi, Marco, et al.
Published: (2023)
When Forgetting Triggers Backdoors: A Clean Unlearning Attack
by: Arazzi, Marco, et al.
Published: (2025)
by: Arazzi, Marco, et al.
Published: (2025)
A Deep Reinforcement Learning Approach for Security-Aware Service Acquisition in IoT
by: Arazzi, Marco, et al.
Published: (2024)
by: Arazzi, Marco, et al.
Published: (2024)
A Novel IoT Trust Model Leveraging Fully Distributed Behavioral Fingerprinting and Secure Delegation
by: Arazzi, Marco, et al.
Published: (2023)
by: Arazzi, Marco, et al.
Published: (2023)
KDk: A Defense Mechanism Against Label Inference Attacks in Vertical Federated Learning
by: Arazzi, Marco, et al.
Published: (2024)
by: Arazzi, Marco, et al.
Published: (2024)
SD-RAG: A Prompt-Injection-Resilient Framework for Selective Disclosure in Retrieval-Augmented Generation
by: Masoud, Aiman Al, et al.
Published: (2026)
by: Masoud, Aiman Al, et al.
Published: (2026)
Privacy Preserving and Robust Aggregation for Cross-Silo Federated Learning in Non-IID Settings
by: Arazzi, Marco, et al.
Published: (2025)
by: Arazzi, Marco, et al.
Published: (2025)
SoK: The Last Line of Defense: On Backdoor Defense Evaluation
by: Abad, Gorka, et al.
Published: (2025)
by: Abad, Gorka, et al.
Published: (2025)
Secure Federated Data Distillation
by: Arazzi, Marco, et al.
Published: (2025)
by: Arazzi, Marco, et al.
Published: (2025)
Subject Data Auditing via Source Inference Attack in Cross-Silo Federated Learning
by: Li, Jiaxin, et al.
Published: (2024)
by: Li, Jiaxin, et al.
Published: (2024)
The SkipSponge Attack: Sponge Weight Poisoning of Deep Neural Networks
by: Lintelo, Jona te, et al.
Published: (2024)
by: Lintelo, Jona te, et al.
Published: (2024)
Enhancing Android Malware Detection with Retrieval-Augmented Generation
by: S., Saraga, et al.
Published: (2025)
by: S., Saraga, et al.
Published: (2025)
Large Language Lobotomy: Jailbreaking Mixture-of-Experts via Expert Silencing
by: Lintelo, Jona te, et al.
Published: (2026)
by: Lintelo, Jona te, et al.
Published: (2026)
BadPatches: Routing-aware Backdoor Attacks on Vision Mixture of Experts
by: Chan, Cedric, et al.
Published: (2025)
by: Chan, Cedric, et al.
Published: (2025)
Membership Privacy Evaluation in Deep Spiking Neural Networks
by: Li, Jiaxin, et al.
Published: (2024)
by: Li, Jiaxin, et al.
Published: (2024)
DroidTTP: Mapping Android Applications with TTP for Cyber Threat Intelligence
by: Arikkat, Dincy R, et al.
Published: (2025)
by: Arikkat, Dincy R, et al.
Published: (2025)
Interpreting Emergent Features in Deep Learning-based Side-channel Analysis
by: Karayalçin, Sengim, et al.
Published: (2025)
by: Karayalçin, Sengim, et al.
Published: (2025)
CatBack: Universal Backdoor Attacks on Tabular Data via Categorical Encoding
by: Tajalli, Behrad, et al.
Published: (2025)
by: Tajalli, Behrad, et al.
Published: (2025)
$$\mathbf{L^2\cdot M = C^2}$$ Large Language Models are Covert Channels
by: Gaure, Simen, et al.
Published: (2024)
by: Gaure, Simen, et al.
Published: (2024)
EmoBack: Backdoor Attacks Against Speaker Identification Using Emotional Prosody
by: Schoof, Coen, et al.
Published: (2024)
by: Schoof, Coen, et al.
Published: (2024)
Context is the Key: Backdoor Attacks for In-Context Learning with Vision Transformers
by: Abad, Gorka, et al.
Published: (2024)
by: Abad, Gorka, et al.
Published: (2024)
Flashy Backdoor: Real-world Environment Backdoor Attack on SNNs with DVS Cameras
by: Riaño, Roberto, et al.
Published: (2024)
by: Riaño, Roberto, et al.
Published: (2024)
NoMod: A Non-modular Attack on Module Learning With Errors
by: Bassotto, Cristian, et al.
Published: (2025)
by: Bassotto, Cristian, et al.
Published: (2025)
Towards Backdoor Stealthiness in Model Parameter Space
by: Xu, Xiaoyun, et al.
Published: (2025)
by: Xu, Xiaoyun, et al.
Published: (2025)
MASCing: Configurable Mixture-of-Experts Behavior via Activation Steering Masks
by: Lintelo, Jona te, et al.
Published: (2026)
by: Lintelo, Jona te, et al.
Published: (2026)
GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
by: Wu, Lichao, et al.
Published: (2025)
by: Wu, Lichao, et al.
Published: (2025)
More is Better (Mostly): On the Backdoor Attacks in Federated Graph Neural Networks
by: Xu, Jing, et al.
Published: (2022)
by: Xu, Jing, et al.
Published: (2022)
Sneaky Spikes: Uncovering Stealthy Backdoor Attacks in Spiking Neural Networks with Neuromorphic Data
by: Abad, Gorka, et al.
Published: (2023)
by: Abad, Gorka, et al.
Published: (2023)
Alleviating the Fear of Losing Alignment in LLM Fine-tuning
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
Privacy-Preserving in Blockchain-based Federated Learning Systems
by: M., Sameera K., et al.
Published: (2024)
by: M., Sameera K., et al.
Published: (2024)
NeuroStrike: Neuron-Level Attacks on Aligned LLMs
by: Wu, Lichao, et al.
Published: (2025)
by: Wu, Lichao, et al.
Published: (2025)
A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
by: Xu, Zihao, et al.
Published: (2024)
by: Xu, Zihao, et al.
Published: (2024)
BAN: Detecting Backdoors Activated by Adversarial Neuron Noise
by: Xu, Xiaoyun, et al.
Published: (2024)
by: Xu, Xiaoyun, et al.
Published: (2024)
Removing the Trigger, Not the Backdoor: Alternative Triggers and Latent Backdoors
by: Abad, Gorka, et al.
Published: (2026)
by: Abad, Gorka, et al.
Published: (2026)
Time-Distributed Backdoor Attacks on Federated Spiking Learning
by: Abad, Gorka, et al.
Published: (2024)
by: Abad, Gorka, et al.
Published: (2024)
Similar Items
-
SecureBreak -- A dataset towards safe and secure models
by: Arazzi, Marco, et al.
Published: (2026) -
Security in LLM-as-a-Judge: A Comprehensive SoK
by: Masoud, Aiman Al, et al.
Published: (2026) -
XBreaking: Understanding how LLMs security alignment can be broken
by: Arazzi, Marco, et al.
Published: (2025) -
Let's Focus: Focused Backdoor Attack against Federated Transfer Learning
by: Arazzi, Marco, et al.
Published: (2024) -
LoRA as Oracle
by: Arazzi, Marco, et al.
Published: (2026)