When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Anonto, Riad Ahmed, Nahiyan, Md Labid Al, Hassan, Md Tanvir |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning
von: Anonto, Riad Ahmed, et al.
Veröffentlicht: (2025)
von: Anonto, Riad Ahmed, et al.
Veröffentlicht: (2025)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025)
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025)
Bengali Text Classification: An Evaluation of Large Language Model Approaches
von: Hoque, Md Mahmudul, et al.
Veröffentlicht: (2026)
von: Hoque, Md Mahmudul, et al.
Veröffentlicht: (2026)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
von: Mamun, Md Abdullah Al, et al.
Veröffentlicht: (2025)
von: Mamun, Md Abdullah Al, et al.
Veröffentlicht: (2025)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
von: Yuan, Youliang, et al.
Veröffentlicht: (2024)
von: Yuan, Youliang, et al.
Veröffentlicht: (2024)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
von: García-Ferrero, Iker, et al.
Veröffentlicht: (2025)
von: García-Ferrero, Iker, et al.
Veröffentlicht: (2025)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
Retrieval Augmented Generation-based Large Language Models for Bridging Transportation Cybersecurity Legal Knowledge Gaps
von: Akbar, Khandakar Ashrafi, et al.
Veröffentlicht: (2025)
von: Akbar, Khandakar Ashrafi, et al.
Veröffentlicht: (2025)
Textual Entailment Recognition with Semantic Features from Empirical Text Representation
von: Shajalal, Md, et al.
Veröffentlicht: (2022)
von: Shajalal, Md, et al.
Veröffentlicht: (2022)
Measuring and Eliminating Refusals in Military Large Language Models
von: FitzGerald, Jack, et al.
Veröffentlicht: (2026)
von: FitzGerald, Jack, et al.
Veröffentlicht: (2026)
What Makes a Good Query? Measuring the Impact of Human-Confusing Linguistic Features on LLM Performance
von: Watson, William, et al.
Veröffentlicht: (2026)
von: Watson, William, et al.
Veröffentlicht: (2026)
Rethinking Word Similarity: Semantic Similarity through Classification Confusion
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2025)
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2025)
A Lightweight Explainable Guardrail for Prompt Safety
von: Islam, Md Asiful, et al.
Veröffentlicht: (2026)
von: Islam, Md Asiful, et al.
Veröffentlicht: (2026)
COSMIC: Generalized Refusal Direction Identification in LLM Activations
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
Identifying and Addressing User-level Security Concerns in Smart Homes Using "Smaller" LLMs
von: Chowdhury, Hafijul Hoque, et al.
Veröffentlicht: (2025)
von: Chowdhury, Hafijul Hoque, et al.
Veröffentlicht: (2025)
ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal
von: Zhang, Haonan, et al.
Veröffentlicht: (2025)
von: Zhang, Haonan, et al.
Veröffentlicht: (2025)
Confusion-Aware Rubric Optimization for LLM-based Automated Grading
von: Chu, Yucheng, et al.
Veröffentlicht: (2026)
von: Chu, Yucheng, et al.
Veröffentlicht: (2026)
LLM-ProS: Analyzing Large Language Models' Performance in Competitive Problem Solving
von: Hossain, Md Sifat, et al.
Veröffentlicht: (2025)
von: Hossain, Md Sifat, et al.
Veröffentlicht: (2025)
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
von: Yuan, Yuan, et al.
Veröffentlicht: (2025)
von: Yuan, Yuan, et al.
Veröffentlicht: (2025)
Sim-CLIP: Unsupervised Siamese Adversarial Fine-Tuning for Robust and Semantically-Rich Vision-Language Models
von: Hossain, Md Zarif, et al.
Veröffentlicht: (2024)
von: Hossain, Md Zarif, et al.
Veröffentlicht: (2024)
The DURel Annotation Tool: Human and Computational Measurement of Semantic Proximity, Sense Clusters and Semantic Change
von: Schlechtweg, Dominik, et al.
Veröffentlicht: (2023)
von: Schlechtweg, Dominik, et al.
Veröffentlicht: (2023)
CompassLLM: A Multi-Agent Approach toward Geo-Spatial Reasoning for Popular Path Query
von: Ananto, Md. Nazmul Islam, et al.
Veröffentlicht: (2025)
von: Ananto, Md. Nazmul Islam, et al.
Veröffentlicht: (2025)
Larger models yield better results? Streamlined severity classification of ADHD-related concerns using BERT-based knowledge distillation
von: Karim, Ahmed Akib Jawad, et al.
Veröffentlicht: (2024)
von: Karim, Ahmed Akib Jawad, et al.
Veröffentlicht: (2024)
FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
von: Zhang, Zhehao, et al.
Veröffentlicht: (2025)
von: Zhang, Zhehao, et al.
Veröffentlicht: (2025)
Semantic Density Effect (SDE): Maximizing Information Per Token Improves LLM Accuracy
von: Ahmed, Amr
Veröffentlicht: (2026)
von: Ahmed, Amr
Veröffentlicht: (2026)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
von: Chae, Kyubyung, et al.
Veröffentlicht: (2025)
von: Chae, Kyubyung, et al.
Veröffentlicht: (2025)
LuxVeri at GenAI Detection Task 3: Cross-Domain Detection of AI-Generated Text Using Inverse Perplexity-Weighted Ensemble of Fine-Tuned Transformer Models
von: Mobin, Md Kamrujjaman, et al.
Veröffentlicht: (2025)
von: Mobin, Md Kamrujjaman, et al.
Veröffentlicht: (2025)
LuxVeri at GenAI Detection Task 1: Inverse Perplexity Weighted Ensemble for Robust Detection of AI-Generated Text across English and Multilingual Contexts
von: Mobin, Md Kamrujjaman, et al.
Veröffentlicht: (2025)
von: Mobin, Md Kamrujjaman, et al.
Veröffentlicht: (2025)
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
von: von Recum, Alexander, et al.
Veröffentlicht: (2024)
von: von Recum, Alexander, et al.
Veröffentlicht: (2024)
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
von: Yang, Junxiao, et al.
Veröffentlicht: (2026)
von: Yang, Junxiao, et al.
Veröffentlicht: (2026)
Xolver: Multi-Agent Reasoning with Holistic Experience Learning Just Like an Olympiad Team
von: Hosain, Md Tanzib, et al.
Veröffentlicht: (2025)
von: Hosain, Md Tanzib, et al.
Veröffentlicht: (2025)
Cross-Lingual Probing and Community-Grounded Analysis of Gender Bias in Low-Resource Bengali
von: Reaj, Md Asgor Hossain, et al.
Veröffentlicht: (2026)
von: Reaj, Md Asgor Hossain, et al.
Veröffentlicht: (2026)
Learn to Refuse: Making Large Language Models More Controllable and Reliable through Knowledge Scope Limitation and Refusal Mechanism
von: Cao, Lang
Veröffentlicht: (2023)
von: Cao, Lang
Veröffentlicht: (2023)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2026)
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2026)
Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2026)
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2026)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
Learn to Disguise: Avoid Refusal Responses in LLM's Defense via a Multi-agent Attacker-Disguiser Game
von: Xu, Qianqiao, et al.
Veröffentlicht: (2024)
von: Xu, Qianqiao, et al.
Veröffentlicht: (2024)
LLMs can be easily Confused by Instructional Distractions
von: Hwang, Yerin, et al.
Veröffentlicht: (2025)
von: Hwang, Yerin, et al.
Veröffentlicht: (2025)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
von: Schwinn, Leo, et al.
Veröffentlicht: (2026)
von: Schwinn, Leo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning
von: Anonto, Riad Ahmed, et al.
Veröffentlicht: (2025) -
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025) -
Bengali Text Classification: An Evaluation of Large Language Model Approaches
von: Hoque, Md Mahmudul, et al.
Veröffentlicht: (2026) -
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
von: Si, Shengyun, et al.
Veröffentlicht: (2025) -
Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
von: Mamun, Md Abdullah Al, et al.
Veröffentlicht: (2025)