XBreaking: Understanding how LLMs security alignment can be broken
Fuente:
arXiv
Saved in:
| Main Authors: | Arazzi, Marco, Kembu, Vignesh Kumar, Nocera, Antonino, P, Vinod |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SecureBreak -- A dataset towards safe and secure models
by: Arazzi, Marco, et al.
Published: (2026)
by: Arazzi, Marco, et al.
Published: (2026)
Privacy Preserving and Robust Aggregation for Cross-Silo Federated Learning in Non-IID Settings
by: Arazzi, Marco, et al.
Published: (2025)
by: Arazzi, Marco, et al.
Published: (2025)
LoRA as Oracle
by: Arazzi, Marco, et al.
Published: (2026)
by: Arazzi, Marco, et al.
Published: (2026)
Security in LLM-as-a-Judge: A Comprehensive SoK
by: Masoud, Aiman Al, et al.
Published: (2026)
by: Masoud, Aiman Al, et al.
Published: (2026)
Privacy-Preserving in Blockchain-based Federated Learning Systems
by: M., Sameera K., et al.
Published: (2024)
by: M., Sameera K., et al.
Published: (2024)
You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation
by: Arazzi, Marco, et al.
Published: (2026)
by: Arazzi, Marco, et al.
Published: (2026)
A Deep Reinforcement Learning Approach for Security-Aware Service Acquisition in IoT
by: Arazzi, Marco, et al.
Published: (2024)
by: Arazzi, Marco, et al.
Published: (2024)
A Novel IoT Trust Model Leveraging Fully Distributed Behavioral Fingerprinting and Secure Delegation
by: Arazzi, Marco, et al.
Published: (2023)
by: Arazzi, Marco, et al.
Published: (2023)
KDk: A Defense Mechanism Against Label Inference Attacks in Vertical Federated Learning
by: Arazzi, Marco, et al.
Published: (2024)
by: Arazzi, Marco, et al.
Published: (2024)
SD-RAG: A Prompt-Injection-Resilient Framework for Selective Disclosure in Retrieval-Augmented Generation
by: Masoud, Aiman Al, et al.
Published: (2026)
by: Masoud, Aiman Al, et al.
Published: (2026)
When Forgetting Triggers Backdoors: A Clean Unlearning Attack
by: Arazzi, Marco, et al.
Published: (2025)
by: Arazzi, Marco, et al.
Published: (2025)
Secure Federated Data Distillation
by: Arazzi, Marco, et al.
Published: (2025)
by: Arazzi, Marco, et al.
Published: (2025)
Subject Data Auditing via Source Inference Attack in Cross-Silo Federated Learning
by: Li, Jiaxin, et al.
Published: (2024)
by: Li, Jiaxin, et al.
Published: (2024)
Let's Focus: Focused Backdoor Attack against Federated Transfer Learning
by: Arazzi, Marco, et al.
Published: (2024)
by: Arazzi, Marco, et al.
Published: (2024)
Label Inference Attacks against Node-level Vertical Federated GNNs
by: Arazzi, Marco, et al.
Published: (2023)
by: Arazzi, Marco, et al.
Published: (2023)
Towards Certified Malware Detection: Provable Guarantees Against Evasion Attacks
by: Giri, Nandakrishna, et al.
Published: (2026)
by: Giri, Nandakrishna, et al.
Published: (2026)
Undetectable Conversations Between AI Agents via Pseudorandom Noise-Resilient Key Exchange
by: Vaikuntanathan, Vinod, et al.
Published: (2026)
by: Vaikuntanathan, Vinod, et al.
Published: (2026)
SoK: The Last Line of Defense: On Backdoor Defense Evaluation
by: Abad, Gorka, et al.
Published: (2025)
by: Abad, Gorka, et al.
Published: (2025)
GShield: Mitigating Poisoning Attacks in Federated Learning
by: M., Sameera K., et al.
Published: (2025)
by: M., Sameera K., et al.
Published: (2025)
Are aligned neural networks adversarially aligned?
by: Carlini, Nicholas, et al.
Published: (2023)
by: Carlini, Nicholas, et al.
Published: (2023)
I can't see it but I can Fine-tune it: On Encrypted Fine-tuning of Transformers using Fully Homomorphic Encryption
by: Panzade, Prajwal, et al.
Published: (2024)
by: Panzade, Prajwal, et al.
Published: (2024)
Sparsity in neural networks can improve their privacy
by: Gonon, Antoine, et al.
Published: (2023)
by: Gonon, Antoine, et al.
Published: (2023)
LLMs can hide text in other text of the same length
by: Norelli, Antonio, et al.
Published: (2025)
by: Norelli, Antonio, et al.
Published: (2025)
Robustness of LLM-enabled vehicle trajectory prediction under data security threats
by: Wang, Feilong, et al.
Published: (2025)
by: Wang, Feilong, et al.
Published: (2025)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
by: Lin, Shi, et al.
Published: (2024)
by: Lin, Shi, et al.
Published: (2024)
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
by: Dubiński, Jan, et al.
Published: (2026)
by: Dubiński, Jan, et al.
Published: (2026)
Automated Consistency Analysis of LLMs
by: Patwardhan, Aditya, et al.
Published: (2025)
by: Patwardhan, Aditya, et al.
Published: (2025)
Large-scale online deanonymization with LLMs
by: Lermen, Simon, et al.
Published: (2026)
by: Lermen, Simon, et al.
Published: (2026)
Scaling Trends for Data Poisoning in LLMs
by: Bowen, Dillon, et al.
Published: (2024)
by: Bowen, Dillon, et al.
Published: (2024)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
by: Chen, Zhuowei, et al.
Published: (2025)
by: Chen, Zhuowei, et al.
Published: (2025)
Leveraging RAG for Training-Free Alignment of LLMs
by: Halloran, John T.
Published: (2026)
by: Halloran, John T.
Published: (2026)
Understanding the Effects of Safety Unalignment on Large Language Models
by: Halloran, John T.
Published: (2026)
by: Halloran, John T.
Published: (2026)
RL-Finetuned LLMs for Privacy-Preserving Synthetic Rewriting
by: Shi, Zhan, et al.
Published: (2025)
by: Shi, Zhan, et al.
Published: (2025)
Agentic Misalignment: How LLMs Could Be Insider Threats
by: Lynch, Aengus, et al.
Published: (2025)
by: Lynch, Aengus, et al.
Published: (2025)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
by: Poppi, Samuele, et al.
Published: (2024)
by: Poppi, Samuele, et al.
Published: (2024)
Fast Exact Unlearning for In-Context Learning Data for LLMs
by: Muresanu, Andrei I., et al.
Published: (2024)
by: Muresanu, Andrei I., et al.
Published: (2024)
Attention Tracker: Detecting Prompt Injection Attacks in LLMs
by: Hung, Kuo-Han, et al.
Published: (2024)
by: Hung, Kuo-Han, et al.
Published: (2024)
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
by: Wang, Erchi, et al.
Published: (2026)
by: Wang, Erchi, et al.
Published: (2026)
Permissioned LLMs: Enforcing Access Control in Large Language Models
by: Jayaraman, Bargav, et al.
Published: (2025)
by: Jayaraman, Bargav, et al.
Published: (2025)
Similar Items
-
SecureBreak -- A dataset towards safe and secure models
by: Arazzi, Marco, et al.
Published: (2026) -
Privacy Preserving and Robust Aggregation for Cross-Silo Federated Learning in Non-IID Settings
by: Arazzi, Marco, et al.
Published: (2025) -
LoRA as Oracle
by: Arazzi, Marco, et al.
Published: (2026) -
Security in LLM-as-a-Judge: A Comprehensive SoK
by: Masoud, Aiman Al, et al.
Published: (2026) -
Privacy-Preserving in Blockchain-based Federated Learning Systems
by: M., Sameera K., et al.
Published: (2024)