Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xinpeng, Hu, Chengzhi, Röttger, Paul, Plank, Barbara |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think
von: Wang, Xinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Xinpeng, et al.
Veröffentlicht: (2024)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
von: Wang, Xinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Xinpeng, et al.
Veröffentlicht: (2024)
Refusal Direction is Universal Across Safety-Aligned Languages
von: Wang, Xinpeng, et al.
Veröffentlicht: (2025)
von: Wang, Xinpeng, et al.
Veröffentlicht: (2025)
No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
von: Plaza-del-Arco, Flor Miriam, et al.
Veröffentlicht: (2025)
von: Plaza-del-Arco, Flor Miriam, et al.
Veröffentlicht: (2025)
Surgical Refusal Ablation: Disentangling Safety from Intelligence via Concept-Guided Spectral Cleaning
von: Cristofano, Tony
Veröffentlicht: (2026)
von: Cristofano, Tony
Veröffentlicht: (2026)
Applying Refusal-Vector Ablation to Llama 3.1 70B Agents
von: Lermen, Simon, et al.
Veröffentlicht: (2024)
von: Lermen, Simon, et al.
Veröffentlicht: (2024)
Analyzing Bias in False Refusal Behavior of Large Language Models for Hate Speech Detoxification
von: Im, Kyuri, et al.
Veröffentlicht: (2026)
von: Im, Kyuri, et al.
Veröffentlicht: (2026)
The Potential and Challenges of Evaluating Attitudes, Opinions, and Values in Large Language Models
von: Ma, Bolei, et al.
Veröffentlicht: (2024)
von: Ma, Bolei, et al.
Veröffentlicht: (2024)
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
von: Zhang, Zhehao, et al.
Veröffentlicht: (2025)
von: Zhang, Zhehao, et al.
Veröffentlicht: (2025)
Understanding When Tree of Thoughts Succeeds: Larger Models Excel in Generation, Not Discrimination
von: Chen, Qiqi, et al.
Veröffentlicht: (2024)
von: Chen, Qiqi, et al.
Veröffentlicht: (2024)
Languages in Whisper-Style Speech Encoders Align Both Phonetically and Semantically
von: Shim, Ryan Soh-Eun, et al.
Veröffentlicht: (2025)
von: Shim, Ryan Soh-Eun, et al.
Veröffentlicht: (2025)
Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models
von: Pernisi, Fabio, et al.
Veröffentlicht: (2024)
von: Pernisi, Fabio, et al.
Veröffentlicht: (2024)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
von: An, Bang, et al.
Veröffentlicht: (2024)
von: An, Bang, et al.
Veröffentlicht: (2024)
Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
von: Frank, Gregory N.
Veröffentlicht: (2026)
von: Frank, Gregory N.
Veröffentlicht: (2026)
Measuring and Mitigating Persona Distortions from AI Writing Assistance
von: Röttger, Paul, et al.
Veröffentlicht: (2026)
von: Röttger, Paul, et al.
Veröffentlicht: (2026)
Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
von: Wang, Xinpeng, et al.
Veröffentlicht: (2025)
von: Wang, Xinpeng, et al.
Veröffentlicht: (2025)
There Is More to Refusal in Large Language Models than a Single Direction
von: Joad, Faaiz, et al.
Veröffentlicht: (2026)
von: Joad, Faaiz, et al.
Veröffentlicht: (2026)
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
Algorithmic Fidelity of Large Language Models in Generating Synthetic German Public Opinions: A Case Study
von: Ma, Bolei, et al.
Veröffentlicht: (2024)
von: Ma, Bolei, et al.
Veröffentlicht: (2024)
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
Please refuse to answer me! Mitigating Over-Refusal in Large Language Models via Adaptive Contrastive Decoding
von: Qi, Yupeng, et al.
Veröffentlicht: (2026)
von: Qi, Yupeng, et al.
Veröffentlicht: (2026)
"Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?
von: Chen, Beiduo, et al.
Veröffentlicht: (2024)
von: Chen, Beiduo, et al.
Veröffentlicht: (2024)
Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2025)
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2025)
If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Models
von: Orth, Jasmin, et al.
Veröffentlicht: (2025)
von: Orth, Jasmin, et al.
Veröffentlicht: (2025)
Refusal in Language Models Is Mediated by a Single Direction
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ
von: Holtermann, Carolin, et al.
Veröffentlicht: (2024)
von: Holtermann, Carolin, et al.
Veröffentlicht: (2024)
DualEdit: Mitigating Safety Fallback in LLM Backdoor Editing via Affirmation-Refusal Regulation
von: Jiang, Houcheng, et al.
Veröffentlicht: (2025)
von: Jiang, Houcheng, et al.
Veröffentlicht: (2025)
Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint
von: Du, Yanrui, et al.
Veröffentlicht: (2025)
von: Du, Yanrui, et al.
Veröffentlicht: (2025)
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
von: Russo, Giuseppe, et al.
Veröffentlicht: (2025)
von: Russo, Giuseppe, et al.
Veröffentlicht: (2025)
LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models
von: Rabern, Brian, et al.
Veröffentlicht: (2026)
von: Rabern, Brian, et al.
Veröffentlicht: (2026)
Understanding Refusal in Language Models with Sparse Autoencoders
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
von: Röttger, Paul, et al.
Veröffentlicht: (2024)
von: Röttger, Paul, et al.
Veröffentlicht: (2024)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
von: Jain, Neel, et al.
Veröffentlicht: (2024)
von: Jain, Neel, et al.
Veröffentlicht: (2024)
Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models
von: Madaan, Lovish, et al.
Veröffentlicht: (2024)
von: Madaan, Lovish, et al.
Veröffentlicht: (2024)
Indirect Question Answering in English, German and Bavarian: A Challenging Task for High- and Low-Resource Languages Alike
von: Winkler, Miriam, et al.
Veröffentlicht: (2026)
von: Winkler, Miriam, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think
von: Wang, Xinpeng, et al.
Veröffentlicht: (2024) -
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
von: Si, Shengyun, et al.
Veröffentlicht: (2025) -
"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
von: Wang, Xinpeng, et al.
Veröffentlicht: (2024) -
Refusal Direction is Universal Across Safety-Aligned Languages
von: Wang, Xinpeng, et al.
Veröffentlicht: (2025) -
No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
von: Plaza-del-Arco, Flor Miriam, et al.
Veröffentlicht: (2025)