Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Campbell, David, Kale, Neil, Sehwag, Udari Madhushani, Herring, Bert, Price, Nick, Borges, Dan, Levinson, Alex, Knight, Christina Q |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
par: Sehwag, Udari Madhushani, et autres
Publié: (2025)
par: Sehwag, Udari Madhushani, et autres
Publié: (2025)
AdvBDGen: Adversarially Fortified Prompt-Specific Fuzzy Backdoor Generator Against LLM Alignment
par: Pathmanathan, Pankayaraj, et autres
Publié: (2024)
par: Pathmanathan, Pankayaraj, et autres
Publié: (2024)
In-Context Learning with Topological Information for Knowledge Graph Completion
par: Sehwag, Udari Madhushani, et autres
Publié: (2024)
par: Sehwag, Udari Madhushani, et autres
Publié: (2024)
ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
par: Lee, Michael S., et autres
Publié: (2026)
par: Lee, Michael S., et autres
Publié: (2026)
GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
par: Xu, Yuancheng, et autres
Publié: (2024)
par: Xu, Yuancheng, et autres
Publié: (2024)
Can LLMs be Scammed? A Baseline Measurement Study
par: Sehwag, Udari Madhushani, et autres
Publié: (2024)
par: Sehwag, Udari Madhushani, et autres
Publié: (2024)
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
par: Xie, Tinghao, et autres
Publié: (2024)
par: Xie, Tinghao, et autres
Publié: (2024)
ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents
par: Sehwag, Udari Madhushani, et autres
Publié: (2026)
par: Sehwag, Udari Madhushani, et autres
Publié: (2026)
Continual Learning of Domain Knowledge from Human Feedback in Text-to-SQL
par: Cook, Thomas, et autres
Publié: (2025)
par: Cook, Thomas, et autres
Publié: (2025)
Collab: Controlled Decoding using Mixture of Agents for LLM Alignment
par: Chakraborty, Souradip, et autres
Publié: (2025)
par: Chakraborty, Souradip, et autres
Publié: (2025)
Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
par: Yin, Qingyu, et autres
Publié: (2025)
par: Yin, Qingyu, et autres
Publié: (2025)
AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration
par: Karthikeyan, Harish, et autres
Publié: (2025)
par: Karthikeyan, Harish, et autres
Publié: (2025)
LHAW: Controllable Underspecification for Long-Horizon Tasks
par: Pu, George, et autres
Publié: (2026)
par: Pu, George, et autres
Publié: (2026)
Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
par: Frank, Gregory N.
Publié: (2026)
par: Frank, Gregory N.
Publié: (2026)
Large Language Models are Autonomous Cyber Defenders
par: Castro, Sebastián R., et autres
Publié: (2025)
par: Castro, Sebastián R., et autres
Publié: (2025)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
par: Hadeliya, Tsimur, et autres
Publié: (2025)
par: Hadeliya, Tsimur, et autres
Publié: (2025)
Uplifted Attackers, Human Defenders: The Cyber Offense-Defense Balance for Trailing-Edge Organizations
par: Murphy, Benjamin, et autres
Publié: (2025)
par: Murphy, Benjamin, et autres
Publié: (2025)
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
par: Wei, Boyi, et autres
Publié: (2025)
par: Wei, Boyi, et autres
Publié: (2025)
The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness
par: Varshney, Neeraj, et autres
Publié: (2023)
par: Varshney, Neeraj, et autres
Publié: (2023)
Deactivating Refusal Triggers: Understanding and Mitigating Overrefusal in Safety Alignment
par: Xue, Zhiyu, et autres
Publié: (2026)
par: Xue, Zhiyu, et autres
Publié: (2026)
Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
par: Mamun, Md Abdullah Al, et autres
Publié: (2025)
par: Mamun, Md Abdullah Al, et autres
Publié: (2025)
Asymmetry by Design: Boosting Cyber Defenders with Differential Access to AI
par: Ee, Shaun, et autres
Publié: (2025)
par: Ee, Shaun, et autres
Publié: (2025)
O3D: Offline Data-driven Discovery and Distillation for Sequential Decision-Making with Large Language Models
par: Xiao, Yuchen, et autres
Publié: (2023)
par: Xiao, Yuchen, et autres
Publié: (2023)
A Case Study on the Use of Representativeness Bias as a Defense Against Adversarial Cyber Threats
par: Hitaj, Briland, et autres
Publié: (2025)
par: Hitaj, Briland, et autres
Publié: (2025)
Oyster-I: Beyond Refusal -- Constructive Safety Alignment for Responsible Language Models
par: Duan, Ranjie, et autres
Publié: (2025)
par: Duan, Ranjie, et autres
Publié: (2025)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
par: Chae, Kyubyung, et autres
Publié: (2025)
par: Chae, Kyubyung, et autres
Publié: (2025)
CyberAlly: Leveraging LLMs and Knowledge Graphs to Empower Cyber Defenders
par: Kim, Minjune, et autres
Publié: (2025)
par: Kim, Minjune, et autres
Publié: (2025)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
par: Chiu, Yu Ying, et autres
Publié: (2025)
par: Chiu, Yu Ying, et autres
Publié: (2025)
Toward a Logic of Generalization about Visualization as a Decision Aid
par: Kale, Alex
Publié: (2025)
par: Kale, Alex
Publié: (2025)
A Red Teaming Roadmap Towards System-Level Safety
par: Wang, Zifan, et autres
Publié: (2025)
par: Wang, Zifan, et autres
Publié: (2025)
Reliable Weak-to-Strong Monitoring of LLM Agents
par: Kale, Neil, et autres
Publié: (2025)
par: Kale, Neil, et autres
Publié: (2025)
The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail?
par: Bhatt, Manish, et autres
Publié: (2026)
par: Bhatt, Manish, et autres
Publié: (2026)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
par: Xiao, Yuxin, et autres
Publié: (2025)
par: Xiao, Yuxin, et autres
Publié: (2025)
When Embedding-Based Defenses Fail: Rethinking Safety in LLM-Based Multi-Agent Systems
par: Zhang, Lingxi, et autres
Publié: (2026)
par: Zhang, Lingxi, et autres
Publié: (2026)
Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
par: Xie, Yuanbo, et autres
Publié: (2025)
par: Xie, Yuanbo, et autres
Publié: (2025)
The Path To Autonomous Cyber Defense
par: Oesch, Sean, et autres
Publié: (2024)
par: Oesch, Sean, et autres
Publié: (2024)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
par: Si, Shengyun, et autres
Publié: (2025)
par: Si, Shengyun, et autres
Publié: (2025)
FedDefender: Backdoor Attack Defense in Federated Learning
par: Gill, Waris, et autres
Publié: (2023)
par: Gill, Waris, et autres
Publié: (2023)
MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment
par: Halloran, John
Publié: (2025)
par: Halloran, John
Publié: (2025)
Characterizing Selective Refusal Bias in Large Language Models
par: Khorramrouz, Adel, et autres
Publié: (2025)
par: Khorramrouz, Adel, et autres
Publié: (2025)
Documents similaires
-
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
par: Sehwag, Udari Madhushani, et autres
Publié: (2025) -
AdvBDGen: Adversarially Fortified Prompt-Specific Fuzzy Backdoor Generator Against LLM Alignment
par: Pathmanathan, Pankayaraj, et autres
Publié: (2024) -
In-Context Learning with Topological Information for Knowledge Graph Completion
par: Sehwag, Udari Madhushani, et autres
Publié: (2024) -
ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
par: Lee, Michael S., et autres
Publié: (2026) -
GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
par: Xu, Yuancheng, et autres
Publié: (2024)