Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Eric Hanchen, Ou, Weixuan, Liu, Run, Pang, Shengyuan, Wan, Guancheng, Duan, Ranjie, Dong, Wei, Chang, Kai-Wei, Wang, XiaoFeng, Wu, Ying Nian, Li, Xinfeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamic Generation of Multi-LLM Agents Communication Topologies with Graph Diffusion Models
by: Jiang, Eric Hanchen, et al.
Published: (2025)
by: Jiang, Eric Hanchen, et al.
Published: (2025)
Learning to Rank Chain-of-Thought: Using a Small Model
by: Jiang, Eric Hanchen, et al.
Published: (2025)
by: Jiang, Eric Hanchen, et al.
Published: (2025)
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
by: Dabas, Mahavir, et al.
Published: (2025)
by: Dabas, Mahavir, et al.
Published: (2025)
Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs
by: Maskey, Utsav, et al.
Published: (2026)
by: Maskey, Utsav, et al.
Published: (2026)
Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
by: Yuan, Shuzhou, et al.
Published: (2025)
by: Yuan, Shuzhou, et al.
Published: (2025)
CompAlign: Improving Compositional Text-to-Image Generation with a Complex Benchmark and Fine-Grained Feedback
by: Wan, Yixin, et al.
Published: (2025)
by: Wan, Yixin, et al.
Published: (2025)
The Male CEO and the Female Assistant: Evaluation and Mitigation of Gender Biases in Text-To-Image Generation of Dual Subjects
by: Wan, Yixin, et al.
Published: (2024)
by: Wan, Yixin, et al.
Published: (2024)
White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs
by: Wan, Yixin, et al.
Published: (2024)
by: Wan, Yixin, et al.
Published: (2024)
OR-Bench: An Over-Refusal Benchmark for Large Language Models
by: Cui, Justin, et al.
Published: (2024)
by: Cui, Justin, et al.
Published: (2024)
Improving Adversarial Robust Fairness via Anti-Bias Soft Label Distillation
by: Zhao, Shiji, et al.
Published: (2023)
by: Zhao, Shiji, et al.
Published: (2023)
Diagnose, Localize, Align: A Full-Stack Framework for Reliable LLM Multi-Agent Systems under Instruction Conflicts
by: Wan, Guancheng, et al.
Published: (2025)
by: Wan, Guancheng, et al.
Published: (2025)
Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors
by: Wang, Weixuan, et al.
Published: (2024)
by: Wang, Weixuan, et al.
Published: (2024)
NDM: A Noise-driven Detection and Mitigation Framework against Implicit Sexual Intentions in Text-to-Image Generation
by: Sun, Yitong, et al.
Published: (2025)
by: Sun, Yitong, et al.
Published: (2025)
InsideOut: Measuring and Mitigating Insider-Outsider Bias in Interview Script Generation
by: Wan, Yixin, et al.
Published: (2025)
by: Wan, Yixin, et al.
Published: (2025)
SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering
by: Maskey, Utsav, et al.
Published: (2025)
by: Maskey, Utsav, et al.
Published: (2025)
Oyster-I: Beyond Refusal -- Constructive Safety Alignment for Responsible Language Models
by: Duan, Ranjie, et al.
Published: (2025)
by: Duan, Ranjie, et al.
Published: (2025)
EVOREFUSE: Evolutionary Prompt Optimization for Evaluation and Mitigation of LLM Over-Refusal to Pseudo-Malicious Instructions
by: Wu, Xiaorui, et al.
Published: (2025)
by: Wu, Xiaorui, et al.
Published: (2025)
Towards Class-wise Fair Adversarial Training via Anti-Bias Soft Label Distillation
by: Zhao, Shiji, et al.
Published: (2025)
by: Zhao, Shiji, et al.
Published: (2025)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
by: Pan, Wenbo, et al.
Published: (2025)
by: Pan, Wenbo, et al.
Published: (2025)
Unlocking the Potential of Text-to-Image Diffusion with PAC-Bayesian Theory
by: Jiang, Eric Hanchen, et al.
Published: (2024)
by: Jiang, Eric Hanchen, et al.
Published: (2024)
ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal
by: Zhang, Haonan, et al.
Published: (2025)
by: Zhang, Haonan, et al.
Published: (2025)
Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning
by: Wang, Shengyuan, et al.
Published: (2025)
by: Wang, Shengyuan, et al.
Published: (2025)
Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations
by: Collu, Matteo Gioele, et al.
Published: (2026)
by: Collu, Matteo Gioele, et al.
Published: (2026)
Stable Coresets via Posterior Sampling: Aligning Induced and Full Loss Landscapes
by: Chang, Wei-Kai, et al.
Published: (2025)
by: Chang, Wei-Kai, et al.
Published: (2025)
Programming Refusal with Conditional Activation Steering
by: Lee, Bruce W., et al.
Published: (2024)
by: Lee, Bruce W., et al.
Published: (2024)
Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
by: Teng, Ma, et al.
Published: (2024)
by: Teng, Ma, et al.
Published: (2024)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
by: Si, Shengyun, et al.
Published: (2025)
by: Si, Shengyun, et al.
Published: (2025)
Please refuse to answer me! Mitigating Over-Refusal in Large Language Models via Adaptive Contrastive Decoding
by: Qi, Yupeng, et al.
Published: (2026)
by: Qi, Yupeng, et al.
Published: (2026)
FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
by: Zhang, Zhehao, et al.
Published: (2025)
by: Zhang, Zhehao, et al.
Published: (2025)
Linearly Decoding Refused Knowledge in Aligned Language Models
by: Shrivastava, Aryan, et al.
Published: (2025)
by: Shrivastava, Aryan, et al.
Published: (2025)
Refusal Direction is Universal Across Safety-Aligned Languages
by: Wang, Xinpeng, et al.
Published: (2025)
by: Wang, Xinpeng, et al.
Published: (2025)
SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators
by: Pruksachatkun, Yada, et al.
Published: (2026)
by: Pruksachatkun, Yada, et al.
Published: (2026)
LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries
by: Ren, Xuancheng, et al.
Published: (2026)
by: Ren, Xuancheng, et al.
Published: (2026)
Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint
by: Du, Yanrui, et al.
Published: (2025)
by: Du, Yanrui, et al.
Published: (2025)
A Vision for Access Control in LLM-based Agent Systems
by: Li, Xinfeng, et al.
Published: (2025)
by: Li, Xinfeng, et al.
Published: (2025)
Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation
by: Cheng, Ruoxi, et al.
Published: (2026)
by: Cheng, Ruoxi, et al.
Published: (2026)
Beyond No: Quantifying AI Over-Refusal and Emotional Attachment Boundaries
by: Noever, David, et al.
Published: (2025)
by: Noever, David, et al.
Published: (2025)
Legilimens: Practical and Unified Content Moderation for Large Language Model Services
by: Wu, Jialin, et al.
Published: (2024)
by: Wu, Jialin, et al.
Published: (2024)
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
by: Liu, Zhenhua, et al.
Published: (2024)
by: Liu, Zhenhua, et al.
Published: (2024)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
by: Hu, Xulin, et al.
Published: (2026)
by: Hu, Xulin, et al.
Published: (2026)
Similar Items
-
Dynamic Generation of Multi-LLM Agents Communication Topologies with Graph Diffusion Models
by: Jiang, Eric Hanchen, et al.
Published: (2025) -
Learning to Rank Chain-of-Thought: Using a Small Model
by: Jiang, Eric Hanchen, et al.
Published: (2025) -
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
by: Dabas, Mahavir, et al.
Published: (2025) -
Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs
by: Maskey, Utsav, et al.
Published: (2026) -
Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
by: Yuan, Shuzhou, et al.
Published: (2025)