Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | von Recum, Alexander, Schnabl, Christoph, Hollbeck, Gabor, Alberti, Silas, Blinde, Philip, von Hagen, Marvin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
by: Si, Shengyun, et al.
Published: (2025)
by: Si, Shengyun, et al.
Published: (2025)
Are Reasoning LLMs Robust to Interventions on Their Chain-of-Thought?
by: von Recum, Alexander, et al.
Published: (2026)
by: von Recum, Alexander, et al.
Published: (2026)
State-Dependent Refusal and Learned Incapacity in RLHF-Aligned Language Models
by: Lee, TK
Published: (2025)
by: Lee, TK
Published: (2025)
Refusal in LLMs is an Affine Function
by: Marshall, Thomas, et al.
Published: (2024)
by: Marshall, Thomas, et al.
Published: (2024)
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts
by: Weidener, Lukas, et al.
Published: (2026)
by: Weidener, Lukas, et al.
Published: (2026)
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
by: Alagharu, Rishab, et al.
Published: (2026)
by: Alagharu, Rishab, et al.
Published: (2026)
Should LLM Safety Be More Than Refusing Harmful Instructions?
by: Maskey, Utsav, et al.
Published: (2025)
by: Maskey, Utsav, et al.
Published: (2025)
LLMs Encode Harmfulness and Refusal Separately
by: Zhao, Jiachen, et al.
Published: (2025)
by: Zhao, Jiachen, et al.
Published: (2025)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
by: Yuan, Youliang, et al.
Published: (2024)
by: Yuan, Youliang, et al.
Published: (2024)
Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
by: Yuan, Shuzhou, et al.
Published: (2025)
by: Yuan, Shuzhou, et al.
Published: (2025)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
by: Pan, Wenbo, et al.
Published: (2025)
by: Pan, Wenbo, et al.
Published: (2025)
Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs
by: Maskey, Utsav, et al.
Published: (2026)
by: Maskey, Utsav, et al.
Published: (2026)
Refusal Speech
by: Nhatuve, Diocleciano, et al.
Published: (2021)
by: Nhatuve, Diocleciano, et al.
Published: (2021)
Refusals of noncitizenship
by: Peter Nyers
Published: (2024)
by: Peter Nyers
Published: (2024)
LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries
by: Ren, Xuancheng, et al.
Published: (2026)
by: Ren, Xuancheng, et al.
Published: (2026)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
Signs of the Great Refusal
by: Siegel, Tedd
Published: (2023)
by: Siegel, Tedd
Published: (2023)
Refusal Dark Matter
by: Christopher Love
Published: (2026)
by: Christopher Love
Published: (2026)
Lithuanian Refusals and Politeness
by: Donata Katinaitė
Published: (2021)
by: Donata Katinaitė
Published: (2021)
Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
by: Kumar, Priyanshu, et al.
Published: (2024)
by: Kumar, Priyanshu, et al.
Published: (2024)
Silenced Biases: The Dark Side LLMs Learned to Refuse
by: Himelstein, Rom, et al.
Published: (2025)
by: Himelstein, Rom, et al.
Published: (2025)
Does Refusal Training in LLMs Generalize to the Past Tense?
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
RAID: Refusal-Aware and Integrated Decoding for Jailbreaking LLMs
by: Nguyen, Tuan T., et al.
Published: (2025)
by: Nguyen, Tuan T., et al.
Published: (2025)
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
by: Liu, Zhenhua, et al.
Published: (2024)
by: Liu, Zhenhua, et al.
Published: (2024)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
by: Hu, Xulin, et al.
Published: (2026)
by: Hu, Xulin, et al.
Published: (2026)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
by: Jain, Neel, et al.
Published: (2024)
by: Jain, Neel, et al.
Published: (2024)
Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations
by: Collu, Matteo Gioele, et al.
Published: (2026)
by: Collu, Matteo Gioele, et al.
Published: (2026)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
by: García-Ferrero, Iker, et al.
Published: (2025)
by: García-Ferrero, Iker, et al.
Published: (2025)
"Garbage" In, "Refuse and Refuse Disposal" Out: Making the Most of the Subject Authority File in OPAC.
by: Horn, Marguerite E.
Published: (2002)
by: Horn, Marguerite E.
Published: (2002)
Refusal Behavior in Large Language Models: A Nonlinear Perspective
by: Hildebrandt, Fabian, et al.
Published: (2025)
by: Hildebrandt, Fabian, et al.
Published: (2025)
RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
by: Asif, Sadia, et al.
Published: (2026)
by: Asif, Sadia, et al.
Published: (2026)
Do Reasoning LLMs Refuse What They Infer in Long Contexts?
by: Fu, Yu, et al.
Published: (2026)
by: Fu, Yu, et al.
Published: (2026)
SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs
by: Ghosh, Shaona, et al.
Published: (2025)
by: Ghosh, Shaona, et al.
Published: (2025)
Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules
by: Pattison, Cameron, et al.
Published: (2026)
by: Pattison, Cameron, et al.
Published: (2026)
To Answer or to Refuse? Investigating the Effect of Refusal to Answer Privacy‐Invasive Question on Applicants' Perceived Hireability
by: Wanlu Li, et al.
Published: (2024)
by: Wanlu Li, et al.
Published: (2024)
“Sorry, I Cannot Fulfill That Request”: Analyzing Large Language Model Responses, Redirections, and Refusals to Polarized News Topics
by: Haley Triem, et al.
Published: (2025)
by: Haley Triem, et al.
Published: (2025)
Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding
by: Lee, Jin-Seop, et al.
Published: (2025)
by: Lee, Jin-Seop, et al.
Published: (2025)
Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal Boundaries
by: Yu, Jiahao, et al.
Published: (2024)
by: Yu, Jiahao, et al.
Published: (2024)
LLMs Can Unlearn Refusal with Only 1,000 Benign Samples
by: Guo, Yangyang, et al.
Published: (2026)
by: Guo, Yangyang, et al.
Published: (2026)
Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
by: Mamun, Md Abdullah Al, et al.
Published: (2025)
by: Mamun, Md Abdullah Al, et al.
Published: (2025)
Similar Items
-
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
by: Si, Shengyun, et al.
Published: (2025) -
Are Reasoning LLMs Robust to Interventions on Their Chain-of-Thought?
by: von Recum, Alexander, et al.
Published: (2026) -
State-Dependent Refusal and Learned Incapacity in RLHF-Aligned Language Models
by: Lee, TK
Published: (2025) -
Refusal in LLMs is an Affine Function
by: Marshall, Thomas, et al.
Published: (2024) -
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts
by: Weidener, Lukas, et al.
Published: (2026)