Universal Refusal Circuits Across LLMs: Cross-Model Transfer via Trajectory Replay and Concept-Basis Reconstruction
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Cristofano, Tony |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Surgical Refusal Ablation: Disentangling Safety from Intelligence via Concept-Guided Spectral Cleaning
par: Cristofano, Tony
Publié: (2026)
par: Cristofano, Tony
Publié: (2026)
Refusal Direction is Universal Across Safety-Aligned Languages
par: Wang, Xinpeng, et autres
Publié: (2025)
par: Wang, Xinpeng, et autres
Publié: (2025)
Self and Cross-Model Distillation for LLMs: Effective Methods for Refusal Pattern Alignment
par: Li, Jie, et autres
Publié: (2024)
par: Li, Jie, et autres
Publié: (2024)
Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer
par: Shao, Shun, et autres
Publié: (2026)
par: Shao, Shun, et autres
Publié: (2026)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
par: Yuan, Youliang, et autres
Publié: (2024)
par: Yuan, Youliang, et autres
Publié: (2024)
LLMs Encode Harmfulness and Refusal Separately
par: Zhao, Jiachen, et autres
Publié: (2025)
par: Zhao, Jiachen, et autres
Publié: (2025)
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
par: Siu, Vincent, et autres
Publié: (2025)
par: Siu, Vincent, et autres
Publié: (2025)
Refusal in LLMs is an Affine Function
par: Marshall, Thomas, et autres
Publié: (2024)
par: Marshall, Thomas, et autres
Publié: (2024)
Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs
par: Maskey, Utsav, et autres
Publié: (2026)
par: Maskey, Utsav, et autres
Publié: (2026)
Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
par: Yuan, Shuzhou, et autres
Publié: (2025)
par: Yuan, Shuzhou, et autres
Publié: (2025)
AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials
par: Xu, Yiheng, et autres
Publié: (2024)
par: Xu, Yiheng, et autres
Publié: (2024)
Benchmarking Concept-Spilling Across Languages in LLMs
par: Badanin, Ilia, et autres
Publié: (2026)
par: Badanin, Ilia, et autres
Publié: (2026)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
par: Si, Shengyun, et autres
Publié: (2025)
par: Si, Shengyun, et autres
Publié: (2025)
RAID: Refusal-Aware and Integrated Decoding for Jailbreaking LLMs
par: Nguyen, Tuan T., et autres
Publié: (2025)
par: Nguyen, Tuan T., et autres
Publié: (2025)
$C$-$ΔΘ$: Circuit-Restricted Weight Arithmetic for Selective Refusal
par: Kasliwal, Aditya, et autres
Publié: (2026)
par: Kasliwal, Aditya, et autres
Publié: (2026)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
par: Hu, Xulin, et autres
Publié: (2026)
par: Hu, Xulin, et autres
Publié: (2026)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
par: Pan, Wenbo, et autres
Publié: (2025)
par: Pan, Wenbo, et autres
Publié: (2025)
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
par: Wollschläger, Tom, et autres
Publié: (2025)
par: Wollschläger, Tom, et autres
Publié: (2025)
Language Model Circuits Are Sparse in the Neuron Basis
par: Arora, Aryaman, et autres
Publié: (2026)
par: Arora, Aryaman, et autres
Publié: (2026)
Silenced Biases: The Dark Side LLMs Learned to Refuse
par: Himelstein, Rom, et autres
Publié: (2025)
par: Himelstein, Rom, et autres
Publié: (2025)
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
par: Liu, Zhenhua, et autres
Publié: (2024)
par: Liu, Zhenhua, et autres
Publié: (2024)
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
par: von Recum, Alexander, et autres
Publié: (2024)
par: von Recum, Alexander, et autres
Publié: (2024)
Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models
par: Yoon, Eunseop, et autres
Publié: (2025)
par: Yoon, Eunseop, et autres
Publié: (2025)
Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
par: Song, Maojia, et autres
Publié: (2024)
par: Song, Maojia, et autres
Publié: (2024)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
par: Jain, Neel, et autres
Publié: (2024)
par: Jain, Neel, et autres
Publié: (2024)
Do Reasoning LLMs Refuse What They Infer in Long Contexts?
par: Fu, Yu, et autres
Publié: (2026)
par: Fu, Yu, et autres
Publié: (2026)
Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts
par: Huang, Youcheng, et autres
Publié: (2025)
par: Huang, Youcheng, et autres
Publié: (2025)
AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling
par: Ding, Liang
Publié: (2026)
par: Ding, Liang
Publié: (2026)
Writer-R1: Enhancing Generative Writing in LLMs via Memory-augmented Replay Policy Optimization
par: Zhao, Jihao, et autres
Publié: (2026)
par: Zhao, Jihao, et autres
Publié: (2026)
Understanding Refusal in Language Models with Sparse Autoencoders
par: Yeo, Wei Jie, et autres
Publié: (2025)
par: Yeo, Wei Jie, et autres
Publié: (2025)
Contrastive Cross-Course Knowledge Tracing via Concept Graph Guided Knowledge Transfer
par: Han, Wenkang, et autres
Publié: (2025)
par: Han, Wenkang, et autres
Publié: (2025)
Transferring Expert Cognitive Models to Social Robots via Agentic Concept Bottleneck Models
par: Zhao, Xinyu, et autres
Publié: (2025)
par: Zhao, Xinyu, et autres
Publié: (2025)
Does Refusal Training in LLMs Generalize to the Past Tense?
par: Andriushchenko, Maksym, et autres
Publié: (2024)
par: Andriushchenko, Maksym, et autres
Publié: (2024)
SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering
par: Maskey, Utsav, et autres
Publié: (2025)
par: Maskey, Utsav, et autres
Publié: (2025)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
par: Muhamed, Aashiq, et autres
Publié: (2025)
par: Muhamed, Aashiq, et autres
Publié: (2025)
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
par: Chhabra, Vishnu Kabir, et autres
Publié: (2025)
par: Chhabra, Vishnu Kabir, et autres
Publié: (2025)
Characterizing Selective Refusal Bias in Large Language Models
par: Khorramrouz, Adel, et autres
Publié: (2025)
par: Khorramrouz, Adel, et autres
Publié: (2025)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
par: Han, Seungju, et autres
Publié: (2024)
par: Han, Seungju, et autres
Publié: (2024)
Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation
par: Wang, Xinpeng, et autres
Publié: (2024)
par: Wang, Xinpeng, et autres
Publié: (2024)
FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
par: Zhang, Zhehao, et autres
Publié: (2025)
par: Zhang, Zhehao, et autres
Publié: (2025)
Documents similaires
-
Surgical Refusal Ablation: Disentangling Safety from Intelligence via Concept-Guided Spectral Cleaning
par: Cristofano, Tony
Publié: (2026) -
Refusal Direction is Universal Across Safety-Aligned Languages
par: Wang, Xinpeng, et autres
Publié: (2025) -
Self and Cross-Model Distillation for LLMs: Effective Methods for Refusal Pattern Alignment
par: Li, Jie, et autres
Publié: (2024) -
Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer
par: Shao, Shun, et autres
Publié: (2026) -
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
par: Yuan, Youliang, et autres
Publié: (2024)