Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Jain, Neel, Shrivastava, Aditya, Zhu, Chenyang, Liu, Daben, Samuel, Alfy, Panda, Ashwinee, Kumar, Anoop, Goldblum, Micah, Goldstein, Tom |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Token Prediction via Self-Distillation
by: Kirchenbauer, John, et al.
Published: (2026)
by: Kirchenbauer, John, et al.
Published: (2026)
LLM Optimization Unlocks Real-Time Pairwise Reranking
by: Wu, Jingyu, et al.
Published: (2025)
by: Wu, Jingyu, et al.
Published: (2025)
FB-RAG: Improving RAG with Forward and Backward Lookup
by: Chawla, Kushal, et al.
Published: (2025)
by: Chawla, Kushal, et al.
Published: (2025)
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
by: Alagharu, Rishab, et al.
Published: (2026)
by: Alagharu, Rishab, et al.
Published: (2026)
A Comparison of Independent and Joint Fine-tuning Strategies for Retrieval-Augmented Generation
by: Lawton, Neal Gregory, et al.
Published: (2025)
by: Lawton, Neal Gregory, et al.
Published: (2025)
Confidence-Based Response Abstinence: Improving LLM Trustworthiness via Activation-Based Uncertainty Estimation
by: Huang, Zhiqi, et al.
Published: (2025)
by: Huang, Zhiqi, et al.
Published: (2025)
Improving Consistency in Retrieval-Augmented Systems with Group Similarity Rewards
by: Hamman, Faisal, et al.
Published: (2025)
by: Hamman, Faisal, et al.
Published: (2025)
Readability Reconsidered: A Cross-Dataset Analysis of Reference-Free Metrics
by: Belem, Catarina G, et al.
Published: (2025)
by: Belem, Catarina G, et al.
Published: (2025)
Linearly Decoding Refused Knowledge in Aligned Language Models
by: Shrivastava, Aryan, et al.
Published: (2025)
by: Shrivastava, Aryan, et al.
Published: (2025)
Play by the Type Rules: Inferring Constraints for LLM Functions in Declarative Programs
by: Glenn, Parker, et al.
Published: (2025)
by: Glenn, Parker, et al.
Published: (2025)
Modeling and Predicting Multi-Turn Answer Instability in Large Language Models
by: He, Jiahang, et al.
Published: (2025)
by: He, Jiahang, et al.
Published: (2025)
FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
by: Hayes, Kevin David, et al.
Published: (2025)
by: Hayes, Kevin David, et al.
Published: (2025)
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
by: Hu, Mengxuan, et al.
Published: (2026)
by: Hu, Mengxuan, et al.
Published: (2026)
Refusal in Language Models Is Mediated by a Single Direction
by: Arditi, Andy, et al.
Published: (2024)
by: Arditi, Andy, et al.
Published: (2024)
LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation
by: Zhang, Juzheng, et al.
Published: (2025)
by: Zhang, Juzheng, et al.
Published: (2025)
$C$-$ΔΘ$: Circuit-Restricted Weight Arithmetic for Selective Refusal
by: Kasliwal, Aditya, et al.
Published: (2026)
by: Kasliwal, Aditya, et al.
Published: (2026)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
by: Wollschläger, Tom, et al.
Published: (2025)
by: Wollschläger, Tom, et al.
Published: (2025)
Characterizing Selective Refusal Bias in Large Language Models
by: Khorramrouz, Adel, et al.
Published: (2025)
by: Khorramrouz, Adel, et al.
Published: (2025)
OR-Bench: An Over-Refusal Benchmark for Large Language Models
by: Cui, Justin, et al.
Published: (2024)
by: Cui, Justin, et al.
Published: (2024)
Measuring and Eliminating Refusals in Military Large Language Models
by: FitzGerald, Jack, et al.
Published: (2026)
by: FitzGerald, Jack, et al.
Published: (2026)
Learn to Refuse: Making Large Language Models More Controllable and Reliable through Knowledge Scope Limitation and Refusal Mechanism
by: Cao, Lang
Published: (2023)
by: Cao, Lang
Published: (2023)
Identifying and Evaluating Inactive Heads in Pretrained LLMs
by: Sandoval-Segura, Pedro, et al.
Published: (2025)
by: Sandoval-Segura, Pedro, et al.
Published: (2025)
Refusal Behavior in Large Language Models: A Nonlinear Perspective
by: Hildebrandt, Fabian, et al.
Published: (2025)
by: Hildebrandt, Fabian, et al.
Published: (2025)
Understanding Refusal in Language Models with Sparse Autoencoders
by: Yeo, Wei Jie, et al.
Published: (2025)
by: Yeo, Wei Jie, et al.
Published: (2025)
DynaGuard: A Dynamic Guardian Model With User-Defined Policies
by: Hoover, Monte, et al.
Published: (2025)
by: Hoover, Monte, et al.
Published: (2025)
There Is More to Refusal in Large Language Models than a Single Direction
by: Joad, Faaiz, et al.
Published: (2026)
by: Joad, Faaiz, et al.
Published: (2026)
Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-Answering
by: Bakman, Yavuz, et al.
Published: (2025)
by: Bakman, Yavuz, et al.
Published: (2025)
Trusted Uncertainty in Large Language Models: A Unified Framework for Confidence Calibration and Risk-Controlled Refusal
by: Oehri, Markus, et al.
Published: (2025)
by: Oehri, Markus, et al.
Published: (2025)
Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
by: Yuan, Shuzhou, et al.
Published: (2025)
by: Yuan, Shuzhou, et al.
Published: (2025)
Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs
by: Maskey, Utsav, et al.
Published: (2026)
by: Maskey, Utsav, et al.
Published: (2026)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
by: Si, Shengyun, et al.
Published: (2025)
by: Si, Shengyun, et al.
Published: (2025)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
by: García-Ferrero, Iker, et al.
Published: (2025)
by: García-Ferrero, Iker, et al.
Published: (2025)
Refusal Direction is Universal Across Safety-Aligned Languages
by: Wang, Xinpeng, et al.
Published: (2025)
by: Wang, Xinpeng, et al.
Published: (2025)
Role-Conditioned Refusals: Evaluating Access Control Reasoning in Large Language Models
by: Klisura, Đorđe, et al.
Published: (2025)
by: Klisura, Đorđe, et al.
Published: (2025)
Refusal in LLMs is an Affine Function
by: Marshall, Thomas, et al.
Published: (2024)
by: Marshall, Thomas, et al.
Published: (2024)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
by: Hu, Xulin, et al.
Published: (2026)
by: Hu, Xulin, et al.
Published: (2026)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
by: Yuan, Youliang, et al.
Published: (2024)
by: Yuan, Youliang, et al.
Published: (2024)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
by: Pan, Wenbo, et al.
Published: (2025)
by: Pan, Wenbo, et al.
Published: (2025)
LLMs Encode Harmfulness and Refusal Separately
by: Zhao, Jiachen, et al.
Published: (2025)
by: Zhao, Jiachen, et al.
Published: (2025)
Similar Items
-
Multi-Token Prediction via Self-Distillation
by: Kirchenbauer, John, et al.
Published: (2026) -
LLM Optimization Unlocks Real-Time Pairwise Reranking
by: Wu, Jingyu, et al.
Published: (2025) -
FB-RAG: Improving RAG with Forward and Backward Lookup
by: Chawla, Kushal, et al.
Published: (2025) -
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
by: Alagharu, Rishab, et al.
Published: (2026) -
A Comparison of Independent and Joint Fine-tuning Strategies for Retrieval-Augmented Generation
by: Lawton, Neal Gregory, et al.
Published: (2025)