Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ferrand, Jean-Charles Noirot, Beugin, Yohan, Pauley, Eric, Sheatsley, Ryan, McDaniel, Patrick |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning
by: Domico, Kyle, et al.
Published: (2025)
by: Domico, Kyle, et al.
Published: (2025)
Characterizing the Modification Space of Signature IDS Rules
by: Guide, Ryan, et al.
Published: (2024)
by: Guide, Ryan, et al.
Published: (2024)
Longitudinal Analyses of SAST Tools: A CodeQL Case Study
by: Ferrand, Jean-Charles Noirot, et al.
Published: (2026)
by: Ferrand, Jean-Charles Noirot, et al.
Published: (2026)
Efficient Storage Integrity in Adversarial Settings
by: Burke, Quinn, et al.
Published: (2025)
by: Burke, Quinn, et al.
Published: (2025)
Secure IP Address Allocation at Cloud Scale
by: Pauley, Eric, et al.
Published: (2022)
by: Pauley, Eric, et al.
Published: (2022)
ParTEETor: A System for Partial Deployments of TEEs within Tor
by: King, Rachel, et al.
Published: (2024)
by: King, Rachel, et al.
Published: (2024)
LibIHT: A Hardware-Based Approach to Efficient and Evasion-Resistant Dynamic Binary Analysis
by: Zhao, Changyu, et al.
Published: (2025)
by: Zhao, Changyu, et al.
Published: (2025)
A Public and Reproducible Assessment of the Topics API on Real Data
by: Beugin, Yohan, et al.
Published: (2024)
by: Beugin, Yohan, et al.
Published: (2024)
Technical Report: The Need for a (Research) Sandstorm through the Privacy Sandbox
by: Beugin, Yohan, et al.
Published: (2025)
by: Beugin, Yohan, et al.
Published: (2025)
The Role of Learning in Attacking ML-based Network Intrusion Detection
by: Domico, Kyle, et al.
Published: (2026)
by: Domico, Kyle, et al.
Published: (2026)
Securing Cloud File Systems with Trusted Execution
by: Burke, Quinn, et al.
Published: (2023)
by: Burke, Quinn, et al.
Published: (2023)
On the Robustness Tradeoff in Fine-Tuning
by: Li, Kunyang, et al.
Published: (2025)
by: Li, Kunyang, et al.
Published: (2025)
Err on the Side of Texture: Texture Bias on Real Data
by: Hoak, Blaine, et al.
Published: (2024)
by: Hoak, Blaine, et al.
Published: (2024)
VisuoAlign: Safety Alignment of LVLMs with Multimodal Tree Search
by: Li, MingSheng, et al.
Published: (2025)
by: Li, MingSheng, et al.
Published: (2025)
On Scalable Integrity Checking for Secure Cloud Disks
by: Burke, Quinn, et al.
Published: (2024)
by: Burke, Quinn, et al.
Published: (2024)
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
by: Wu, Fangzhou, et al.
Published: (2024)
by: Wu, Fangzhou, et al.
Published: (2024)
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense
by: Zhang, Jiawen, et al.
Published: (2025)
by: Zhang, Jiawen, et al.
Published: (2025)
MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
by: Guo, Weiyang, et al.
Published: (2025)
by: Guo, Weiyang, et al.
Published: (2025)
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment
by: Kim, Jaehan, et al.
Published: (2025)
by: Kim, Jaehan, et al.
Published: (2025)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Reimagining Safety Alignment with An Image
by: Xia, Yifan, et al.
Published: (2025)
by: Xia, Yifan, et al.
Published: (2025)
FreakOut-LLM: The Effect of Emotional Stimuli on Safety Alignment
by: Kuznetsov, Daniel, et al.
Published: (2026)
by: Kuznetsov, Daniel, et al.
Published: (2026)
Multi-Stream Perturbation Attack: Breaking Safety Alignment of Thinking LLMs Through Concurrent Task Interference
by: Yang, Fan
Published: (2026)
by: Yang, Fan
Published: (2026)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off
by: Chen, Yu, et al.
Published: (2026)
by: Chen, Yu, et al.
Published: (2026)
AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
by: Liu, Xiaogeng, et al.
Published: (2024)
by: Liu, Xiaogeng, et al.
Published: (2024)
Agent Safety Alignment via Reinforcement Learning
by: Sha, Zeyang, et al.
Published: (2025)
by: Sha, Zeyang, et al.
Published: (2025)
Safety Layers in Aligned Large Language Models: The Key to LLM Security
by: Li, Shen, et al.
Published: (2024)
by: Li, Shen, et al.
Published: (2024)
AgentAlign: Navigating Safety Alignment in the Shift from Informative to Agentic Large Language Models
by: Zhang, Jinchuan, et al.
Published: (2025)
by: Zhang, Jinchuan, et al.
Published: (2025)
Measuring Safety Alignment Effects in Autonomous Security Agents
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
Robustifying Safety-Aligned Large Language Models through Clean Data Curation
by: Liu, Xiaoqun, et al.
Published: (2024)
by: Liu, Xiaoqun, et al.
Published: (2024)
Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models
by: Luo, Weidi, et al.
Published: (2025)
by: Luo, Weidi, et al.
Published: (2025)
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
Matching Ranks Over Probability Yields Truly Deep Safety Alignment
by: Vega, Jason, et al.
Published: (2025)
by: Vega, Jason, et al.
Published: (2025)
Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
by: Campbell, David, et al.
Published: (2026)
by: Campbell, David, et al.
Published: (2026)
PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
by: Li, Nanxi, et al.
Published: (2025)
by: Li, Nanxi, et al.
Published: (2025)
MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation
by: Yan, Lu, et al.
Published: (2025)
by: Yan, Lu, et al.
Published: (2025)
On The Dangers of Poisoned LLMs In Security Automation
by: Karlsen, Patrick, et al.
Published: (2025)
by: Karlsen, Patrick, et al.
Published: (2025)
Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
by: Yin, Qingyu, et al.
Published: (2025)
by: Yin, Qingyu, et al.
Published: (2025)
Similar Items
-
Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning
by: Domico, Kyle, et al.
Published: (2025) -
Characterizing the Modification Space of Signature IDS Rules
by: Guide, Ryan, et al.
Published: (2024) -
Longitudinal Analyses of SAST Tools: A CodeQL Case Study
by: Ferrand, Jean-Charles Noirot, et al.
Published: (2026) -
Efficient Storage Integrity in Adversarial Settings
by: Burke, Quinn, et al.
Published: (2025) -
Secure IP Address Allocation at Cloud Scale
by: Pauley, Eric, et al.
Published: (2022)