SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering
Fuente:
arXiv
Salvato in:
| Autori principali: | Maskey, Utsav, Yadav, Sumit, Dras, Mark, Naseem, Usman |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs
di: Maskey, Utsav, et al.
Pubblicazione: (2026)
di: Maskey, Utsav, et al.
Pubblicazione: (2026)
Steering Over-refusals Towards Safety in Retrieval Augmented Generation
di: Maskey, Utsav, et al.
Pubblicazione: (2025)
di: Maskey, Utsav, et al.
Pubblicazione: (2025)
Should LLM Safety Be More Than Refusing Harmful Instructions?
di: Maskey, Utsav, et al.
Pubblicazione: (2025)
di: Maskey, Utsav, et al.
Pubblicazione: (2025)
Steering Towards Fairness: Mitigating Political Bias in LLMs
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
Fairness Evaluation and Inference Level Mitigation in LLMs
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilities
di: Maskey, Utsav, et al.
Pubblicazione: (2025)
di: Maskey, Utsav, et al.
Pubblicazione: (2025)
MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language
di: Yadav, Sumit, et al.
Pubblicazione: (2025)
di: Yadav, Sumit, et al.
Pubblicazione: (2025)
Framing Political Bias in Multilingual LLMs Across Pakistani Languages
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2025)
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2025)
SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs
di: Ren, Juan, et al.
Pubblicazione: (2025)
di: Ren, Juan, et al.
Pubblicazione: (2025)
Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack
di: Ren, Juan, et al.
Pubblicazione: (2025)
di: Ren, Juan, et al.
Pubblicazione: (2025)
DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models
di: Ren, Kaixuan, et al.
Pubblicazione: (2025)
di: Ren, Kaixuan, et al.
Pubblicazione: (2025)
Too Helpful, Too Harmless, Too Honest or Just Right?
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2025)
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2025)
AlignCultura: Towards Culturally Aligned Large Language Models?
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2026)
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2026)
When the Model Said 'No Comment', We Knew Helpfulness Was Dead, Honesty Was Alive, and Safety Was Terrified
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2026)
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2026)
Beyond the Black Box: Demystifying Multi-Turn LLM Reasoning with VISTA
di: Zhang, Yiran, et al.
Pubblicazione: (2025)
di: Zhang, Yiran, et al.
Pubblicazione: (2025)
CogMem: A Cognitive Memory Architecture for Sustained Multi-Turn Reasoning in Large Language Models
di: Zhang, Yiran, et al.
Pubblicazione: (2025)
di: Zhang, Yiran, et al.
Pubblicazione: (2025)
VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare
di: Shetty, Anudeex, et al.
Pubblicazione: (2025)
di: Shetty, Anudeex, et al.
Pubblicazione: (2025)
Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
di: Yuan, Shuzhou, et al.
Pubblicazione: (2025)
di: Yuan, Shuzhou, et al.
Pubblicazione: (2025)
Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs
di: Nadeem, Afrozah, et al.
Pubblicazione: (2026)
di: Nadeem, Afrozah, et al.
Pubblicazione: (2026)
A Survey on Progress in LLM Alignment from the Perspective of Reward Design
di: Ji, Miaomiao, et al.
Pubblicazione: (2025)
di: Ji, Miaomiao, et al.
Pubblicazione: (2025)
Activation-Space Personality Steering: Hybrid Layer Selection for Stable Trait Control in LLMs
di: Bhandari, Pranav, et al.
Pubblicazione: (2025)
di: Bhandari, Pranav, et al.
Pubblicazione: (2025)
VaxGuard: A Multi-Generator, Multi-Type, and Multi-Role Dataset for Detecting LLM-Generated Vaccine Misinformation
di: Ahmad, Syed Talal, et al.
Pubblicazione: (2025)
di: Ahmad, Syed Talal, et al.
Pubblicazione: (2025)
Mechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions
di: Naseem, Usman
Pubblicazione: (2026)
di: Naseem, Usman
Pubblicazione: (2026)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
di: Pan, Wenbo, et al.
Pubblicazione: (2025)
di: Pan, Wenbo, et al.
Pubblicazione: (2025)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
di: Cheng, Stephen, et al.
Pubblicazione: (2026)
di: Cheng, Stephen, et al.
Pubblicazione: (2026)
Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models
di: Bhandari, Pranav, et al.
Pubblicazione: (2026)
di: Bhandari, Pranav, et al.
Pubblicazione: (2026)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
di: Si, Shengyun, et al.
Pubblicazione: (2025)
di: Si, Shengyun, et al.
Pubblicazione: (2025)
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
di: Liu, Zhenhua, et al.
Pubblicazione: (2024)
di: Liu, Zhenhua, et al.
Pubblicazione: (2024)
FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
di: Zhang, Zhehao, et al.
Pubblicazione: (2025)
di: Zhang, Zhehao, et al.
Pubblicazione: (2025)
RAID: Refusal-Aware and Integrated Decoding for Jailbreaking LLMs
di: Nguyen, Tuan T., et al.
Pubblicazione: (2025)
di: Nguyen, Tuan T., et al.
Pubblicazione: (2025)
Can Reasoning LLMs Enhance Clinical Document Classification?
di: Mustafa, Akram, et al.
Pubblicazione: (2025)
di: Mustafa, Akram, et al.
Pubblicazione: (2025)
Evaluating Hierarchical Clinical Document Classification Using Reasoning-Based LLMs
di: Mustafa, Akram, et al.
Pubblicazione: (2025)
di: Mustafa, Akram, et al.
Pubblicazione: (2025)
Mitigating Memorization in LLMs using Activation Steering
di: Suri, Manan, et al.
Pubblicazione: (2025)
di: Suri, Manan, et al.
Pubblicazione: (2025)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
di: García-Ferrero, Iker, et al.
Pubblicazione: (2025)
di: García-Ferrero, Iker, et al.
Pubblicazione: (2025)
Programming Refusal with Conditional Activation Steering
di: Lee, Bruce W., et al.
Pubblicazione: (2024)
di: Lee, Bruce W., et al.
Pubblicazione: (2024)
GRAIT: Gradient-Driven Refusal-Aware Instruction Tuning for Effective Hallucination Mitigation
di: Zhu, Runchuan, et al.
Pubblicazione: (2025)
di: Zhu, Runchuan, et al.
Pubblicazione: (2025)
Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models
di: Ren, Juan, et al.
Pubblicazione: (2025)
di: Ren, Juan, et al.
Pubblicazione: (2025)
Flick: Few Labels Text Classification using K-Aware Intermediate Learning in Multi-Task Low-Resource Languages
di: Almutairi, Ali, et al.
Pubblicazione: (2025)
di: Almutairi, Ali, et al.
Pubblicazione: (2025)
Myanmar XNLI: Building a Dataset and Exploring Low-resource Approaches to Natural Language Inference with Myanmar
di: Htet, Aung Kyaw, et al.
Pubblicazione: (2025)
di: Htet, Aung Kyaw, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs
di: Maskey, Utsav, et al.
Pubblicazione: (2026) -
Steering Over-refusals Towards Safety in Retrieval Augmented Generation
di: Maskey, Utsav, et al.
Pubblicazione: (2025) -
Should LLM Safety Be More Than Refusing Harmful Instructions?
di: Maskey, Utsav, et al.
Pubblicazione: (2025) -
Steering Towards Fairness: Mitigating Political Bias in LLMs
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025) -
Fairness Evaluation and Inference Level Mitigation in LLMs
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)