Saved in:
| Main Author: | Haralambiev, Kristiyan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.25861 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating False Alarm and Missing Attacks in CAN IDS
by: Hossain, Nirab, et al.
Published: (2026)
by: Hossain, Nirab, et al.
Published: (2026)
Probing the Robustness of Large Language Models Safety to Latent Perturbations
by: Gu, Tianle, et al.
Published: (2025)
by: Gu, Tianle, et al.
Published: (2025)
Mind the Gap: Missing Cyber Threat Coverage in NIDS Datasets for the Energy Sector
by: Tory, Adrita Rahman, et al.
Published: (2025)
by: Tory, Adrita Rahman, et al.
Published: (2025)
From Data Leak to Secret Misses: The Impact of Data Leakage on Secret Detection Models
by: Soltaniani, Farnaz, et al.
Published: (2026)
by: Soltaniani, Farnaz, et al.
Published: (2026)
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
by: Huang, Tiansheng, et al.
Published: (2025)
by: Huang, Tiansheng, et al.
Published: (2025)
When Intelligence Fails: An Empirical Study on Why LLMs Struggle with Password Cracking
by: Rehman, Mohammad Abdul, et al.
Published: (2025)
by: Rehman, Mohammad Abdul, et al.
Published: (2025)
Saffron-1: Safety Inference Scaling
by: Qiu, Ruizhong, et al.
Published: (2025)
by: Qiu, Ruizhong, et al.
Published: (2025)
Quantized Delta Weight Is Safety Keeper
by: Liu, Yule, et al.
Published: (2024)
by: Liu, Yule, et al.
Published: (2024)
Self-Mined Hardness for Safety Fine-Tuning
by: Gupta, Prakhar, et al.
Published: (2026)
by: Gupta, Prakhar, et al.
Published: (2026)
Fast and Lightweight Backdoor Detection via Head Random Probing
by: Yu, Yinbo, et al.
Published: (2026)
by: Yu, Yinbo, et al.
Published: (2026)
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
by: Zloczower, Itay, et al.
Published: (2026)
by: Zloczower, Itay, et al.
Published: (2026)
Understanding the Effects of Safety Unalignment on Large Language Models
by: Halloran, John T.
Published: (2026)
by: Halloran, John T.
Published: (2026)
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
by: Li, Yuxi, et al.
Published: (2024)
by: Li, Yuxi, et al.
Published: (2024)
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
by: Min, Rui, et al.
Published: (2024)
by: Min, Rui, et al.
Published: (2024)
RASA: Routing-Aware Safety Alignment for Mixture-of-Experts Models
by: Liang, Jiacheng, et al.
Published: (2026)
by: Liang, Jiacheng, et al.
Published: (2026)
GAVEL: Towards Rule-Based Safety Through Activation Monitoring
by: Rozenfeld, Shir, et al.
Published: (2026)
by: Rozenfeld, Shir, et al.
Published: (2026)
A Safety and Security Framework for Real-World Agentic Systems
by: Ghosh, Shaona, et al.
Published: (2025)
by: Ghosh, Shaona, et al.
Published: (2025)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
by: Jiang, Fengqing, et al.
Published: (2025)
by: Jiang, Fengqing, et al.
Published: (2025)
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
by: Zaree, Pedram, et al.
Published: (2025)
by: Zaree, Pedram, et al.
Published: (2025)
Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States
by: Chia, Xin Wei, et al.
Published: (2025)
by: Chia, Xin Wei, et al.
Published: (2025)
Probing Network Decisions: Capturing Uncertainties and Unveiling Vulnerabilities Without Label Information
by: Joung, Youngju, et al.
Published: (2025)
by: Joung, Youngju, et al.
Published: (2025)
VulCatch: Enhancing Binary Vulnerability Detection through CodeT5 Decompilation and KAN Advanced Feature Extraction
by: Chukkol, Abdulrahman Hamman Adama, et al.
Published: (2024)
by: Chukkol, Abdulrahman Hamman Adama, et al.
Published: (2024)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models
by: Fang, Junfeng, et al.
Published: (2025)
by: Fang, Junfeng, et al.
Published: (2025)
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem
by: Lin, Shuyi, et al.
Published: (2025)
by: Lin, Shuyi, et al.
Published: (2025)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models
by: Hossain, Ismail, et al.
Published: (2026)
by: Hossain, Ismail, et al.
Published: (2026)
Knowing without Acting: The Disentangled Geometry of Safety Mechanisms in Large Language Models
by: Wu, Jinman, et al.
Published: (2026)
by: Wu, Jinman, et al.
Published: (2026)
UpSafe$^\circ$C: Upcycling for Controllable Safety in Large Language Models
by: Sun, Yuhao, et al.
Published: (2025)
by: Sun, Yuhao, et al.
Published: (2025)
MCP Safety Audit: LLMs with the Model Context Protocol Allow Major Security Exploits
by: Radosevich, Brandon, et al.
Published: (2025)
by: Radosevich, Brandon, et al.
Published: (2025)
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
by: Lu, Ning, et al.
Published: (2025)
by: Lu, Ning, et al.
Published: (2025)
Secure LLM Fine-Tuning via Safety-Aware Probing
by: Wu, Chengcan, et al.
Published: (2025)
by: Wu, Chengcan, et al.
Published: (2025)
SafeRedirect: Defeating Internal Safety Collapse via Task-Completion Redirection in Frontier LLMs
by: Pan, Chao, et al.
Published: (2026)
by: Pan, Chao, et al.
Published: (2026)
Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
by: Liu, Guozhi, et al.
Published: (2025)
by: Liu, Guozhi, et al.
Published: (2025)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms
by: Kazdan, Joshua, et al.
Published: (2025)
by: Kazdan, Joshua, et al.
Published: (2025)
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
by: Luo, Zeren, et al.
Published: (2025)
by: Luo, Zeren, et al.
Published: (2025)
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection
by: Chen, Shuhao, et al.
Published: (2026)
by: Chen, Shuhao, et al.
Published: (2026)
Similar Items
-
Evaluating False Alarm and Missing Attacks in CAN IDS
by: Hossain, Nirab, et al.
Published: (2026) -
Probing the Robustness of Large Language Models Safety to Latent Perturbations
by: Gu, Tianle, et al.
Published: (2025) -
Mind the Gap: Missing Cyber Threat Coverage in NIDS Datasets for the Energy Sector
by: Tory, Adrita Rahman, et al.
Published: (2025) -
From Data Leak to Secret Misses: The Impact of Data Leakage on Secret Detection Models
by: Soltaniani, Farnaz, et al.
Published: (2026) -
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
by: Huang, Tiansheng, et al.
Published: (2025)