Should LLM Safety Be More Than Refusing Harmful Instructions?
Fuente:
arXiv
Saved in:
| Main Authors: | Maskey, Utsav, Dras, Mark, Naseem, Usman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs
by: Maskey, Utsav, et al.
Published: (2026)
by: Maskey, Utsav, et al.
Published: (2026)
Steering Over-refusals Towards Safety in Retrieval Augmented Generation
by: Maskey, Utsav, et al.
Published: (2025)
by: Maskey, Utsav, et al.
Published: (2025)
SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering
by: Maskey, Utsav, et al.
Published: (2025)
by: Maskey, Utsav, et al.
Published: (2025)
Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilities
by: Maskey, Utsav, et al.
Published: (2025)
by: Maskey, Utsav, et al.
Published: (2025)
When the Model Said 'No Comment', We Knew Helpfulness Was Dead, Honesty Was Alive, and Safety Was Terrified
by: Kashyap, Gautam Siddharth, et al.
Published: (2026)
by: Kashyap, Gautam Siddharth, et al.
Published: (2026)
SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs
by: Ren, Juan, et al.
Published: (2025)
by: Ren, Juan, et al.
Published: (2025)
Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack
by: Ren, Juan, et al.
Published: (2025)
by: Ren, Juan, et al.
Published: (2025)
Beyond the Black Box: Demystifying Multi-Turn LLM Reasoning with VISTA
by: Zhang, Yiran, et al.
Published: (2025)
by: Zhang, Yiran, et al.
Published: (2025)
Steering Towards Fairness: Mitigating Political Bias in LLMs
by: Nadeem, Afrozah, et al.
Published: (2025)
by: Nadeem, Afrozah, et al.
Published: (2025)
Fairness Evaluation and Inference Level Mitigation in LLMs
by: Nadeem, Afrozah, et al.
Published: (2025)
by: Nadeem, Afrozah, et al.
Published: (2025)
Framing Political Bias in Multilingual LLMs Across Pakistani Languages
by: Nadeem, Afrozah, et al.
Published: (2025)
by: Nadeem, Afrozah, et al.
Published: (2025)
We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
Too Helpful, Too Harmless, Too Honest or Just Right?
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
AlignCultura: Towards Culturally Aligned Large Language Models?
by: Kashyap, Gautam Siddharth, et al.
Published: (2026)
by: Kashyap, Gautam Siddharth, et al.
Published: (2026)
CogMem: A Cognitive Memory Architecture for Sustained Multi-Turn Reasoning in Large Language Models
by: Zhang, Yiran, et al.
Published: (2025)
by: Zhang, Yiran, et al.
Published: (2025)
VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare
by: Shetty, Anudeex, et al.
Published: (2025)
by: Shetty, Anudeex, et al.
Published: (2025)
MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language
by: Yadav, Sumit, et al.
Published: (2025)
by: Yadav, Sumit, et al.
Published: (2025)
DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models
by: Ren, Kaixuan, et al.
Published: (2025)
by: Ren, Kaixuan, et al.
Published: (2025)
A Survey on Progress in LLM Alignment from the Perspective of Reward Design
by: Ji, Miaomiao, et al.
Published: (2025)
by: Ji, Miaomiao, et al.
Published: (2025)
VaxGuard: A Multi-Generator, Multi-Type, and Multi-Role Dataset for Detecting LLM-Generated Vaccine Misinformation
by: Ahmad, Syed Talal, et al.
Published: (2025)
by: Ahmad, Syed Talal, et al.
Published: (2025)
Mechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions
by: Naseem, Usman
Published: (2026)
by: Naseem, Usman
Published: (2026)
Confidence Should Be Calibrated More Than One Turn Deep
by: Zhang, Zhaohan, et al.
Published: (2026)
by: Zhang, Zhaohan, et al.
Published: (2026)
Safety Alignment Should Be Made More Than Just A Few Attention Heads
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
LLMs Encode Harmfulness and Refusal Separately
by: Zhao, Jiachen, et al.
Published: (2025)
by: Zhao, Jiachen, et al.
Published: (2025)
PersoDPO: Scalable Preference Optimization for Instruction-Adherent, Persona-Grounded Dialogue via Multi-LLM Evaluation
by: Afzoon, Saleh, et al.
Published: (2026)
by: Afzoon, Saleh, et al.
Published: (2026)
Safety-Aligned Weights Are Not Enough: Refusal-Teacher-Guided Finetuning Enhances Safety and Downstream Performance under Harmful Finetuning Attacks
by: Ham, Seokil, et al.
Published: (2025)
by: Ham, Seokil, et al.
Published: (2025)
XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content
by: Abishethvarman, Vadivel, et al.
Published: (2025)
by: Abishethvarman, Vadivel, et al.
Published: (2025)
Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models
by: Ren, Juan, et al.
Published: (2025)
by: Ren, Juan, et al.
Published: (2025)
Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity
by: Qu, Jiaming, et al.
Published: (2026)
by: Qu, Jiaming, et al.
Published: (2026)
Myanmar XNLI: Building a Dataset and Exploring Low-resource Approaches to Natural Language Inference with Myanmar
by: Htet, Aung Kyaw, et al.
Published: (2025)
by: Htet, Aung Kyaw, et al.
Published: (2025)
Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety
by: Zhang, Yuyou, et al.
Published: (2025)
by: Zhang, Yuyou, et al.
Published: (2025)
When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
by: Anonto, Riad Ahmed, et al.
Published: (2025)
by: Anonto, Riad Ahmed, et al.
Published: (2025)
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
by: von Recum, Alexander, et al.
Published: (2024)
by: von Recum, Alexander, et al.
Published: (2024)
Enhancing ESG Impact Type Identification through Early Fusion and Multilingual Models
by: Veeramani, Hariram, et al.
Published: (2024)
by: Veeramani, Hariram, et al.
Published: (2024)
PersoPilot: An Adaptive AI-Copilot for Transparent Contextualized Persona Classification and Personalized Response Generation
by: Afzoon, Saleh, et al.
Published: (2026)
by: Afzoon, Saleh, et al.
Published: (2026)
DualEdit: Mitigating Safety Fallback in LLM Backdoor Editing via Affirmation-Refusal Regulation
by: Jiang, Houcheng, et al.
Published: (2025)
by: Jiang, Houcheng, et al.
Published: (2025)
Do Large Language Models Reflect Demographic Pluralism in Safety?
by: Naseem, Usman, et al.
Published: (2026)
by: Naseem, Usman, et al.
Published: (2026)
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026)
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026)
From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered
by: Devic, Siddartha, et al.
Published: (2025)
by: Devic, Siddartha, et al.
Published: (2025)
Refusal Direction is Universal Across Safety-Aligned Languages
by: Wang, Xinpeng, et al.
Published: (2025)
by: Wang, Xinpeng, et al.
Published: (2025)
Similar Items
-
Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs
by: Maskey, Utsav, et al.
Published: (2026) -
Steering Over-refusals Towards Safety in Retrieval Augmented Generation
by: Maskey, Utsav, et al.
Published: (2025) -
SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering
by: Maskey, Utsav, et al.
Published: (2025) -
Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilities
by: Maskey, Utsav, et al.
Published: (2025) -
When the Model Said 'No Comment', We Knew Helpfulness Was Dead, Honesty Was Alive, and Safety Was Terrified
by: Kashyap, Gautam Siddharth, et al.
Published: (2026)