Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuan, Youliang, Jiao, Wenxiang, Wang, Wenxuan, Huang, Jen-tse, Xu, Jiahao, Liang, Tian, He, Pinjia, Tu, Zhaopeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
von: Yuan, Youliang, et al.
Veröffentlicht: (2023)
von: Yuan, Youliang, et al.
Veröffentlicht: (2023)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs
von: Liu, Xiaoyuan, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoyuan, et al.
Veröffentlicht: (2024)
All Languages Matter: On the Multilingual Safety of Large Language Models
von: Wang, Wenxuan, et al.
Veröffentlicht: (2023)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2023)
Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
von: Yuan, Youliang, et al.
Veröffentlicht: (2025)
von: Yuan, Youliang, et al.
Veröffentlicht: (2025)
LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
von: Wan, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wan, Yuxuan, et al.
Veröffentlicht: (2024)
Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench
von: Huang, Jen-tse, et al.
Veröffentlicht: (2023)
von: Huang, Jen-tse, et al.
Veröffentlicht: (2023)
Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by Step
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
On the Shortcut Learning in Multilingual Neural Machine Translation
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
von: Huang, Jen-tse, et al.
Veröffentlicht: (2024)
von: Huang, Jen-tse, et al.
Veröffentlicht: (2024)
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
von: Huang, Jen-Tse, et al.
Veröffentlicht: (2025)
von: Huang, Jen-Tse, et al.
Veröffentlicht: (2025)
Who is ChatGPT? Benchmarking LLMs' Psychological Portrayal Using PsychoBench
von: Huang, Jen-tse, et al.
Veröffentlicht: (2023)
von: Huang, Jen-tse, et al.
Veröffentlicht: (2023)
Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
von: Yuan, Youliang, et al.
Veröffentlicht: (2025)
von: Yuan, Youliang, et al.
Veröffentlicht: (2025)
Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language Models
von: Wang, Wenxuan, et al.
Veröffentlicht: (2023)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2023)
Identifying the Achilles' Heel: An Iterative Method for Dynamically Uncovering Factual Errors in Large Language Models
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMs
von: Zhao, Sihang, et al.
Veröffentlicht: (2026)
von: Zhao, Sihang, et al.
Veröffentlicht: (2026)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
von: Kumar, Priyanshu, et al.
Veröffentlicht: (2024)
von: Kumar, Priyanshu, et al.
Veröffentlicht: (2024)
Latent Adversarial Training Improves the Representation of Refusal
von: Abbas, Alexandra, et al.
Veröffentlicht: (2025)
von: Abbas, Alexandra, et al.
Veröffentlicht: (2025)
Difficult Task Yes but Simple Task No: Unveiling the Laziness in Multimodal LLMs
von: Zhao, Sihang, et al.
Veröffentlicht: (2024)
von: Zhao, Sihang, et al.
Veröffentlicht: (2024)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
von: Liu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoyuan, et al.
Veröffentlicht: (2025)
VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Models
von: Huang, Jen-tse, et al.
Veröffentlicht: (2025)
von: Huang, Jen-tse, et al.
Veröffentlicht: (2025)
Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
LLMs Encode Harmfulness and Refusal Separately
von: Zhao, Jiachen, et al.
Veröffentlicht: (2025)
von: Zhao, Jiachen, et al.
Veröffentlicht: (2025)
Does Refusal Training in LLMs Generalize to the Past Tense?
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
von: Yuan, Yuan, et al.
Veröffentlicht: (2025)
von: Yuan, Yuan, et al.
Veröffentlicht: (2025)
Refusal in LLMs is an Affine Function
von: Marshall, Thomas, et al.
Veröffentlicht: (2024)
von: Marshall, Thomas, et al.
Veröffentlicht: (2024)
RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
von: Asif, Sadia, et al.
Veröffentlicht: (2026)
von: Asif, Sadia, et al.
Veröffentlicht: (2026)
SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs
von: Ghosh, Shaona, et al.
Veröffentlicht: (2025)
von: Ghosh, Shaona, et al.
Veröffentlicht: (2025)
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts
von: Weidener, Lukas, et al.
Veröffentlicht: (2026)
von: Weidener, Lukas, et al.
Veröffentlicht: (2026)
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
von: Alagharu, Rishab, et al.
Veröffentlicht: (2026)
von: Alagharu, Rishab, et al.
Veröffentlicht: (2026)
Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
von: Yin, Qingyu, et al.
Veröffentlicht: (2025)
von: Yin, Qingyu, et al.
Veröffentlicht: (2025)
Understanding and Mitigating the Uncertainty in Zero-Shot Translation
von: Wang, Wenxuan, et al.
Veröffentlicht: (2022)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2022)
Learning to Ask: When LLM Agents Meet Unclear Instruction
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries
von: Ren, Xuancheng, et al.
Veröffentlicht: (2026)
von: Ren, Xuancheng, et al.
Veröffentlicht: (2026)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs
von: Maskey, Utsav, et al.
Veröffentlicht: (2026)
von: Maskey, Utsav, et al.
Veröffentlicht: (2026)
Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal Boundaries
von: Yu, Jiahao, et al.
Veröffentlicht: (2024)
von: Yu, Jiahao, et al.
Veröffentlicht: (2024)
HumorReject: Decoupling LLM Safety from Refusal Prefix via A Little Humor
von: Wu, Zihui, et al.
Veröffentlicht: (2025)
von: Wu, Zihui, et al.
Veröffentlicht: (2025)
MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment
von: Halloran, John
Veröffentlicht: (2025)
von: Halloran, John
Veröffentlicht: (2025)
Ähnliche Einträge
-
GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
von: Yuan, Youliang, et al.
Veröffentlicht: (2023) -
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025) -
Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs
von: Liu, Xiaoyuan, et al.
Veröffentlicht: (2024) -
All Languages Matter: On the Multilingual Safety of Large Language Models
von: Wang, Wenxuan, et al.
Veröffentlicht: (2023) -
Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
von: Yuan, Youliang, et al.
Veröffentlicht: (2025)