Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications
Fuente:
arXiv
Saved in:
| Main Authors: | David, Isaac, Gervais, Arthur |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Measuring Safety Alignment Effects in Autonomous Security Agents
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
Multi-Agent Penetration Testing AI for the Web
by: David, Isaac, et al.
Published: (2025)
by: David, Isaac, et al.
Published: (2025)
AuthorMist: Evading AI Text Detectors with Reinforcement Learning
by: David, Isaac, et al.
Published: (2025)
by: David, Isaac, et al.
Published: (2025)
Towards Optimal Agentic Architectures for Offensive Security Tasks
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
AI Agent Smart Contract Exploit Generation
by: Gervais, Arthur, et al.
Published: (2025)
by: Gervais, Arthur, et al.
Published: (2025)
Alignment Contracts for Agentic Security Systems
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models
by: Zeng, Yi, et al.
Published: (2024)
by: Zeng, Yi, et al.
Published: (2024)
Safety Layers in Aligned Large Language Models: The Key to LLM Security
by: Li, Shen, et al.
Published: (2024)
by: Li, Shen, et al.
Published: (2024)
SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers
by: Qin, Kaihua, et al.
Published: (2026)
by: Qin, Kaihua, et al.
Published: (2026)
Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
by: Xie, Yuanbo, et al.
Published: (2025)
by: Xie, Yuanbo, et al.
Published: (2025)
TxRay: Agentic Postmortem of Live Blockchain Attacks
by: Wang, Ziyue, et al.
Published: (2026)
by: Wang, Ziyue, et al.
Published: (2026)
Reimagining Safety Alignment with An Image
by: Xia, Yifan, et al.
Published: (2025)
by: Xia, Yifan, et al.
Published: (2025)
Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models
by: Tan, Rui Yang, et al.
Published: (2026)
by: Tan, Rui Yang, et al.
Published: (2026)
Systematization of Knowledge: Security and Safety in the Model Context Protocol Ecosystem
by: Gaire, Shiva, et al.
Published: (2025)
by: Gaire, Shiva, et al.
Published: (2025)
Privacy-Preserving Large Language Models: Mechanisms, Applications, and Future Directions
by: Zhao, Guoshenghui, et al.
Published: (2024)
by: Zhao, Guoshenghui, et al.
Published: (2024)
Agent Safety Alignment via Reinforcement Learning
by: Sha, Zeyang, et al.
Published: (2025)
by: Sha, Zeyang, et al.
Published: (2025)
Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
by: Campbell, David, et al.
Published: (2026)
by: Campbell, David, et al.
Published: (2026)
Decompiling Smart Contracts with a Large Language Model
by: David, Isaac, et al.
Published: (2025)
by: David, Isaac, et al.
Published: (2025)
aiXamine: Simplified LLM Safety and Security
by: Deniz, Fatih, et al.
Published: (2025)
by: Deniz, Fatih, et al.
Published: (2025)
(Security) Assertions by Large Language Models
by: Kande, Rahul, et al.
Published: (2023)
by: Kande, Rahul, et al.
Published: (2023)
Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs
by: Ferrand, Jean-Charles Noirot, et al.
Published: (2025)
by: Ferrand, Jean-Charles Noirot, et al.
Published: (2025)
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
by: Du, Pengfei
Published: (2025)
by: Du, Pengfei
Published: (2025)
Lifelong Safety Alignment for Language Models
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Cisco Integrated AI Security and Safety Framework Report
by: Chang, Amy, et al.
Published: (2025)
by: Chang, Amy, et al.
Published: (2025)
SoK: Towards Security and Safety of Edge AI
by: Wingarz, Tatjana, et al.
Published: (2024)
by: Wingarz, Tatjana, et al.
Published: (2024)
Emerging Security Challenges of Large Language Models
by: Debar, Herve, et al.
Published: (2024)
by: Debar, Herve, et al.
Published: (2024)
FreakOut-LLM: The Effect of Emotional Stimuli on Safety Alignment
by: Kuznetsov, Daniel, et al.
Published: (2026)
by: Kuznetsov, Daniel, et al.
Published: (2026)
VisuoAlign: Safety Alignment of LVLMs with Multimodal Tree Search
by: Li, MingSheng, et al.
Published: (2025)
by: Li, MingSheng, et al.
Published: (2025)
Mobile Application Threats and Security
by: Mirzoev, Timur, et al.
Published: (2025)
by: Mirzoev, Timur, et al.
Published: (2025)
AI Risk Management Should Incorporate Both Safety and Security
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
ASTRIDE: A Security Threat Modeling Platform for Agentic-AI Applications
by: Bandara, Eranga, et al.
Published: (2025)
by: Bandara, Eranga, et al.
Published: (2025)
Security Concerns for Large Language Models: A Survey
by: Li, Miles Q., et al.
Published: (2025)
by: Li, Miles Q., et al.
Published: (2025)
A Survey on Data Security in Large Language Models
by: Chen, Kang, et al.
Published: (2025)
by: Chen, Kang, et al.
Published: (2025)
Matching Ranks Over Probability Yields Truly Deep Safety Alignment
by: Vega, Jason, et al.
Published: (2025)
by: Vega, Jason, et al.
Published: (2025)
PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
by: Li, Nanxi, et al.
Published: (2025)
by: Li, Nanxi, et al.
Published: (2025)
PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and Restoration
by: Zeng, Ziqian, et al.
Published: (2024)
by: Zeng, Ziqian, et al.
Published: (2024)
SafeThinker: Reasoning about Risk to Deepen Safety Beyond Shallow Alignment
by: Fang, Xianya, et al.
Published: (2026)
by: Fang, Xianya, et al.
Published: (2026)
Hallucination-Resistant Security Planning with a Large Language Model
by: Hammar, Kim, et al.
Published: (2026)
by: Hammar, Kim, et al.
Published: (2026)
Similar Items
-
Measuring Safety Alignment Effects in Autonomous Security Agents
by: David, Isaac, et al.
Published: (2026) -
Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches
by: David, Isaac, et al.
Published: (2026) -
Multi-Agent Penetration Testing AI for the Web
by: David, Isaac, et al.
Published: (2025) -
AuthorMist: Evading AI Text Detectors with Reinforcement Learning
by: David, Isaac, et al.
Published: (2025) -
Towards Optimal Agentic Architectures for Offensive Security Tasks
by: David, Isaac, et al.
Published: (2026)