COSMIC: Generalized Refusal Direction Identification in LLM Activations
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Siu, Vincent, Crispino, Nicholas, Yu, Zihao, Pan, Sam, Wang, Zhun, Liu, Yang, Song, Dawn, Wang, Chenguang |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
par: Siu, Vincent, et autres
Publié: (2025)
par: Siu, Vincent, et autres
Publié: (2025)
Peer-Preservation in Frontier Models
par: Potter, Yujin, et autres
Publié: (2026)
par: Potter, Yujin, et autres
Publié: (2026)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
par: Siu, Vincent, et autres
Publié: (2025)
par: Siu, Vincent, et autres
Publié: (2025)
Agent Instructs Large Language Models to be General Zero-Shot Reasoners
par: Crispino, Nicholas, et autres
Publié: (2023)
par: Crispino, Nicholas, et autres
Publié: (2023)
A Framework for Formalizing LLM Agent Security
par: Siu, Vincent, et autres
Publié: (2026)
par: Siu, Vincent, et autres
Publié: (2026)
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
par: Wang, Zhun, et autres
Publié: (2025)
par: Wang, Zhun, et autres
Publié: (2025)
LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
par: Kolasani, Sai, et autres
Publié: (2025)
par: Kolasani, Sai, et autres
Publié: (2025)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
par: Wang, Xilong, et autres
Publié: (2026)
par: Wang, Xilong, et autres
Publié: (2026)
Benchmarking Zero-Shot Robustness of Multimodal Foundation Models: A Pilot Study
par: Wang, Chenguang, et autres
Publié: (2024)
par: Wang, Chenguang, et autres
Publié: (2024)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
par: García-Ferrero, Iker, et autres
Publié: (2025)
par: García-Ferrero, Iker, et autres
Publié: (2025)
Re-Tuning: Overcoming the Compositionality Limits of Large Language Models with Recursive Tuning
par: Pasewark, Eric, et autres
Publié: (2024)
par: Pasewark, Eric, et autres
Publié: (2024)
Programming Refusal with Conditional Activation Steering
par: Lee, Bruce W., et autres
Publié: (2024)
par: Lee, Bruce W., et autres
Publié: (2024)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
par: Pan, Wenbo, et autres
Publié: (2025)
par: Pan, Wenbo, et autres
Publié: (2025)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
par: Si, Shengyun, et autres
Publié: (2025)
par: Si, Shengyun, et autres
Publié: (2025)
ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal
par: Zhang, Haonan, et autres
Publié: (2025)
par: Zhang, Haonan, et autres
Publié: (2025)
Learn to Disguise: Avoid Refusal Responses in LLM's Defense via a Multi-agent Attacker-Disguiser Game
par: Xu, Qianqiao, et autres
Publié: (2024)
par: Xu, Qianqiao, et autres
Publié: (2024)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
par: Yuan, Youliang, et autres
Publié: (2024)
par: Yuan, Youliang, et autres
Publié: (2024)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
par: Muhamed, Aashiq, et autres
Publié: (2025)
par: Muhamed, Aashiq, et autres
Publié: (2025)
COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning
par: Pan, Jing, et autres
Publié: (2023)
par: Pan, Jing, et autres
Publié: (2023)
Refusal in Language Models Is Mediated by a Single Direction
par: Arditi, Andy, et autres
Publié: (2024)
par: Arditi, Andy, et autres
Publié: (2024)
Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?
par: Wang, Qineng, et autres
Publié: (2024)
par: Wang, Qineng, et autres
Publié: (2024)
When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
par: Anonto, Riad Ahmed, et autres
Publié: (2025)
par: Anonto, Riad Ahmed, et autres
Publié: (2025)
Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy
par: Jiang, Eric Hanchen, et autres
Publié: (2025)
par: Jiang, Eric Hanchen, et autres
Publié: (2025)
Investigating Bias Representations in Llama 2 Chat via Activation Steering
par: Lu, Dawn, et autres
Publié: (2024)
par: Lu, Dawn, et autres
Publié: (2024)
Generate-on-Graph: Treat LLM as both Agent and KG in Incomplete Knowledge Graph Question Answering
par: Xu, Yao, et autres
Publié: (2024)
par: Xu, Yao, et autres
Publié: (2024)
Predicting Task Performance with Context-aware Scaling Laws
par: Montgomery, Kyle, et autres
Publié: (2025)
par: Montgomery, Kyle, et autres
Publié: (2025)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
par: Hu, Xulin, et autres
Publié: (2026)
par: Hu, Xulin, et autres
Publié: (2026)
$\texttt{COSMIC}$: Mutual Information for Task-Agnostic Summarization Evaluation
par: Darrin, Maxime, et autres
Publié: (2024)
par: Darrin, Maxime, et autres
Publié: (2024)
Measuring and Eliminating Refusals in Military Large Language Models
par: FitzGerald, Jack, et autres
Publié: (2026)
par: FitzGerald, Jack, et autres
Publié: (2026)
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
par: Pan, Wenbo, et autres
Publié: (2025)
par: Pan, Wenbo, et autres
Publié: (2025)
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
par: Liu, Zhenhua, et autres
Publié: (2024)
par: Liu, Zhenhua, et autres
Publié: (2024)
Effectively Steer LLM To Follow Preference via Building Confident Directions
par: Song, Bingqing, et autres
Publié: (2025)
par: Song, Bingqing, et autres
Publié: (2025)
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
par: von Recum, Alexander, et autres
Publié: (2024)
par: von Recum, Alexander, et autres
Publié: (2024)
Activation-Guided Local Editing for Jailbreaking Attacks
par: Wang, Jiecong, et autres
Publié: (2025)
par: Wang, Jiecong, et autres
Publié: (2025)
Does Refusal Training in LLMs Generalize to the Past Tense?
par: Andriushchenko, Maksym, et autres
Publié: (2024)
par: Andriushchenko, Maksym, et autres
Publié: (2024)
Learn to Refuse: Making Large Language Models More Controllable and Reliable through Knowledge Scope Limitation and Refusal Mechanism
par: Cao, Lang
Publié: (2023)
par: Cao, Lang
Publié: (2023)
PlanGenLLMs: A Modern Survey of LLM Planning Capabilities
par: Wei, Hui, et autres
Publié: (2025)
par: Wei, Hui, et autres
Publié: (2025)
dLLM: Simple Diffusion Language Modeling
par: Zhou, Zhanhui, et autres
Publié: (2026)
par: Zhou, Zhanhui, et autres
Publié: (2026)
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
par: Wollschläger, Tom, et autres
Publié: (2025)
par: Wollschläger, Tom, et autres
Publié: (2025)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
par: Hadeliya, Tsimur, et autres
Publié: (2025)
par: Hadeliya, Tsimur, et autres
Publié: (2025)
Documents similaires
-
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
par: Siu, Vincent, et autres
Publié: (2025) -
Peer-Preservation in Frontier Models
par: Potter, Yujin, et autres
Publié: (2026) -
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
par: Siu, Vincent, et autres
Publié: (2025) -
Agent Instructs Large Language Models to be General Zero-Shot Reasoners
par: Crispino, Nicholas, et autres
Publié: (2023) -
A Framework for Formalizing LLM Agent Security
par: Siu, Vincent, et autres
Publié: (2026)