Saved in:
| Main Authors: | Anonto, Riad Ahmed, Nahiyan, Md Labid Al, Hassan, Md Tanvir |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.01037 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning
by: Anonto, Riad Ahmed, et al.
Published: (2025)
by: Anonto, Riad Ahmed, et al.
Published: (2025)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
by: Hadeliya, Tsimur, et al.
Published: (2025)
by: Hadeliya, Tsimur, et al.
Published: (2025)
Bengali Text Classification: An Evaluation of Large Language Model Approaches
by: Hoque, Md Mahmudul, et al.
Published: (2026)
by: Hoque, Md Mahmudul, et al.
Published: (2026)
Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
by: Mamun, Md Abdullah Al, et al.
Published: (2025)
by: Mamun, Md Abdullah Al, et al.
Published: (2025)
Identifying and Addressing User-level Security Concerns in Smart Homes Using "Smaller" LLMs
by: Chowdhury, Hafijul Hoque, et al.
Published: (2025)
by: Chowdhury, Hafijul Hoque, et al.
Published: (2025)
Retrieval Augmented Generation-based Large Language Models for Bridging Transportation Cybersecurity Legal Knowledge Gaps
by: Akbar, Khandakar Ashrafi, et al.
Published: (2025)
by: Akbar, Khandakar Ashrafi, et al.
Published: (2025)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
by: Si, Shengyun, et al.
Published: (2025)
by: Si, Shengyun, et al.
Published: (2025)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
by: Yuan, Youliang, et al.
Published: (2024)
by: Yuan, Youliang, et al.
Published: (2024)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
by: García-Ferrero, Iker, et al.
Published: (2025)
by: García-Ferrero, Iker, et al.
Published: (2025)
Textual Entailment Recognition with Semantic Features from Empirical Text Representation
by: Shajalal, Md, et al.
Published: (2022)
by: Shajalal, Md, et al.
Published: (2022)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
by: Pan, Wenbo, et al.
Published: (2025)
by: Pan, Wenbo, et al.
Published: (2025)
A Lightweight Explainable Guardrail for Prompt Safety
by: Islam, Md Asiful, et al.
Published: (2026)
by: Islam, Md Asiful, et al.
Published: (2026)
Measuring and Eliminating Refusals in Military Large Language Models
by: FitzGerald, Jack, et al.
Published: (2026)
by: FitzGerald, Jack, et al.
Published: (2026)
ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal
by: Zhang, Haonan, et al.
Published: (2025)
by: Zhang, Haonan, et al.
Published: (2025)
Rethinking Word Similarity: Semantic Similarity through Classification Confusion
by: Zhou, Kaitlyn, et al.
Published: (2025)
by: Zhou, Kaitlyn, et al.
Published: (2025)
What Makes a Good Query? Measuring the Impact of Human-Confusing Linguistic Features on LLM Performance
by: Watson, William, et al.
Published: (2026)
by: Watson, William, et al.
Published: (2026)
COSMIC: Generalized Refusal Direction Identification in LLM Activations
by: Siu, Vincent, et al.
Published: (2025)
by: Siu, Vincent, et al.
Published: (2025)
Sim-CLIP: Unsupervised Siamese Adversarial Fine-Tuning for Robust and Semantically-Rich Vision-Language Models
by: Hossain, Md Zarif, et al.
Published: (2024)
by: Hossain, Md Zarif, et al.
Published: (2024)
Confusion-Aware Rubric Optimization for LLM-based Automated Grading
by: Chu, Yucheng, et al.
Published: (2026)
by: Chu, Yucheng, et al.
Published: (2026)
LLM-ProS: Analyzing Large Language Models' Performance in Competitive Problem Solving
by: Hossain, Md Sifat, et al.
Published: (2025)
by: Hossain, Md Sifat, et al.
Published: (2025)
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
by: Yuan, Yuan, et al.
Published: (2025)
by: Yuan, Yuan, et al.
Published: (2025)
Larger models yield better results? Streamlined severity classification of ADHD-related concerns using BERT-based knowledge distillation
by: Karim, Ahmed Akib Jawad, et al.
Published: (2024)
by: Karim, Ahmed Akib Jawad, et al.
Published: (2024)
CompassLLM: A Multi-Agent Approach toward Geo-Spatial Reasoning for Popular Path Query
by: Ananto, Md. Nazmul Islam, et al.
Published: (2025)
by: Ananto, Md. Nazmul Islam, et al.
Published: (2025)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
by: Kabir, Mohsinul, et al.
Published: (2026)
by: Kabir, Mohsinul, et al.
Published: (2026)
Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum
by: Al-Radhi, Mohammed Salah, et al.
Published: (2026)
by: Al-Radhi, Mohammed Salah, et al.
Published: (2026)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
by: Chae, Kyubyung, et al.
Published: (2025)
by: Chae, Kyubyung, et al.
Published: (2025)
The DURel Annotation Tool: Human and Computational Measurement of Semantic Proximity, Sense Clusters and Semantic Change
by: Schlechtweg, Dominik, et al.
Published: (2023)
by: Schlechtweg, Dominik, et al.
Published: (2023)
Cross-Lingual Probing and Community-Grounded Analysis of Gender Bias in Low-Resource Bengali
by: Reaj, Md Asgor Hossain, et al.
Published: (2026)
by: Reaj, Md Asgor Hossain, et al.
Published: (2026)
LuxVeri at GenAI Detection Task 3: Cross-Domain Detection of AI-Generated Text Using Inverse Perplexity-Weighted Ensemble of Fine-Tuned Transformer Models
by: Mobin, Md Kamrujjaman, et al.
Published: (2025)
by: Mobin, Md Kamrujjaman, et al.
Published: (2025)
LuxVeri at GenAI Detection Task 1: Inverse Perplexity Weighted Ensemble for Robust Detection of AI-Generated Text across English and Multilingual Contexts
by: Mobin, Md Kamrujjaman, et al.
Published: (2025)
by: Mobin, Md Kamrujjaman, et al.
Published: (2025)
FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
by: Zhang, Zhehao, et al.
Published: (2025)
by: Zhang, Zhehao, et al.
Published: (2025)
Xolver: Multi-Agent Reasoning with Holistic Experience Learning Just Like an Olympiad Team
by: Hosain, Md Tanzib, et al.
Published: (2025)
by: Hosain, Md Tanzib, et al.
Published: (2025)
Semantic Density Effect (SDE): Maximizing Information Per Token Improves LLM Accuracy
by: Ahmed, Amr
Published: (2026)
by: Ahmed, Amr
Published: (2026)
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
by: Yang, Junxiao, et al.
Published: (2026)
by: Yang, Junxiao, et al.
Published: (2026)
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
by: von Recum, Alexander, et al.
Published: (2024)
by: von Recum, Alexander, et al.
Published: (2024)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
Learn to Refuse: Making Large Language Models More Controllable and Reliable through Knowledge Scope Limitation and Refusal Mechanism
by: Cao, Lang
Published: (2023)
by: Cao, Lang
Published: (2023)
SPARE: Single-Pass Annotation with Reference-Guided Evaluation for Automatic Process Supervision and Reward Modelling
by: Rizvi, Md Imbesat Hassan, et al.
Published: (2025)
by: Rizvi, Md Imbesat Hassan, et al.
Published: (2025)
SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models
by: Rizvi, Md Imbesat Hassan, et al.
Published: (2024)
by: Rizvi, Md Imbesat Hassan, et al.
Published: (2024)
LLMs can be easily Confused by Instructional Distractions
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
Similar Items
-
Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning
by: Anonto, Riad Ahmed, et al.
Published: (2025) -
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
by: Hadeliya, Tsimur, et al.
Published: (2025) -
Bengali Text Classification: An Evaluation of Large Language Model Approaches
by: Hoque, Md Mahmudul, et al.
Published: (2026) -
Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
by: Mamun, Md Abdullah Al, et al.
Published: (2025) -
Identifying and Addressing User-level Security Concerns in Smart Homes Using "Smaller" LLMs
by: Chowdhury, Hafijul Hoque, et al.
Published: (2025)