Refusal in LLMs is an Affine Function
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Marshall, Thomas, Scherlis, Adam, Belrose, Nora |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Understanding Gradient Descent through the Training Jacobian
par: Belrose, Nora, et autres
Publié: (2024)
par: Belrose, Nora, et autres
Publié: (2024)
Estimating the Probability of Sampling a Trained Neural Network at Random
par: Scherlis, Adam, et autres
Publié: (2025)
par: Scherlis, Adam, et autres
Publié: (2025)
Does Transformer Interpretability Transfer to RNNs?
par: Paulo, Gonçalo, et autres
Publié: (2024)
par: Paulo, Gonçalo, et autres
Publié: (2024)
Partially Rewriting a Transformer in Natural Language
par: Paulo, Gonçalo, et autres
Publié: (2025)
par: Paulo, Gonçalo, et autres
Publié: (2025)
Mechanistic Anomaly Detection for "Quirky" Language Models
par: Johnston, David O., et autres
Publié: (2025)
par: Johnston, David O., et autres
Publié: (2025)
Automatically Interpreting Millions of Features in Large Language Models
par: Paulo, Gonçalo, et autres
Publié: (2024)
par: Paulo, Gonçalo, et autres
Publié: (2024)
Eliciting Latent Knowledge from Quirky Language Models
par: Mallen, Alex, et autres
Publié: (2023)
par: Mallen, Alex, et autres
Publié: (2023)
Latent Adversarial Training Improves the Representation of Refusal
par: Abbas, Alexandra, et autres
Publié: (2025)
par: Abbas, Alexandra, et autres
Publié: (2025)
Silenced Biases: The Dark Side LLMs Learned to Refuse
par: Himelstein, Rom, et autres
Publié: (2025)
par: Himelstein, Rom, et autres
Publié: (2025)
LEACE: Perfect linear concept erasure in closed form
par: Belrose, Nora, et autres
Publié: (2023)
par: Belrose, Nora, et autres
Publié: (2023)
Does Refusal Training in LLMs Generalize to the Past Tense?
par: Andriushchenko, Maksym, et autres
Publié: (2024)
par: Andriushchenko, Maksym, et autres
Publié: (2024)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
par: Jain, Neel, et autres
Publié: (2024)
par: Jain, Neel, et autres
Publié: (2024)
LATMiX: Learnable Affine Transformations for Microscaling Quantization of LLMs
par: Gordon, Ofir, et autres
Publié: (2026)
par: Gordon, Ofir, et autres
Publié: (2026)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
par: Muhamed, Aashiq, et autres
Publié: (2025)
par: Muhamed, Aashiq, et autres
Publié: (2025)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
par: Du, Hongzhe, et autres
Publié: (2025)
par: Du, Hongzhe, et autres
Publié: (2025)
Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
par: Mamun, Md Abdullah Al, et autres
Publié: (2025)
par: Mamun, Md Abdullah Al, et autres
Publié: (2025)
Balancing Label Quantity and Quality for Scalable Elicitation
par: Mallen, Alex, et autres
Publié: (2024)
par: Mallen, Alex, et autres
Publié: (2024)
Programming Refusal with Conditional Activation Steering
par: Lee, Bruce W., et autres
Publié: (2024)
par: Lee, Bruce W., et autres
Publié: (2024)
Evaluating SAE interpretability without explanations
par: Paulo, Gonçalo, et autres
Publié: (2025)
par: Paulo, Gonçalo, et autres
Publié: (2025)
Slowing Learning by Erasing Simple Features
par: Quirke, Lucia, et autres
Publié: (2025)
par: Quirke, Lucia, et autres
Publié: (2025)
Sparse Autoencoders Trained on the Same Data Learn Different Features
par: Paulo, Gonçalo, et autres
Publié: (2025)
par: Paulo, Gonçalo, et autres
Publié: (2025)
Converting MLPs into Polynomials in Closed Form
par: Belrose, Nora, et autres
Publié: (2025)
par: Belrose, Nora, et autres
Publié: (2025)
Where Do Reasoning Models Refuse?
par: Yamaguchi, Kureha, et autres
Publié: (2025)
par: Yamaguchi, Kureha, et autres
Publié: (2025)
On Affine Homotopy between Language Encoders
par: Chan, Robin SM, et autres
Publié: (2024)
par: Chan, Robin SM, et autres
Publié: (2024)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
par: Hu, Xulin, et autres
Publié: (2026)
par: Hu, Xulin, et autres
Publié: (2026)
RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
par: Asif, Sadia, et autres
Publié: (2026)
par: Asif, Sadia, et autres
Publié: (2026)
Refusal in Language Models Is Mediated by a Single Direction
par: Arditi, Andy, et autres
Publié: (2024)
par: Arditi, Andy, et autres
Publié: (2024)
Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry
par: Lan, Wenhao, et autres
Publié: (2026)
par: Lan, Wenhao, et autres
Publié: (2026)
Examining Two Hop Reasoning Through Information Content Scaling
par: Johnston, David, et autres
Publié: (2025)
par: Johnston, David, et autres
Publié: (2025)
Representation Surgery: Theory and Practice of Affine Steering
par: Singh, Shashwat, et autres
Publié: (2024)
par: Singh, Shashwat, et autres
Publié: (2024)
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
par: Wollschläger, Tom, et autres
Publié: (2025)
par: Wollschläger, Tom, et autres
Publié: (2025)
Geometric Properties of the Voronoi Tessellation in Latent Semantic Manifolds of Large Language Models
par: Brett, Marshall
Publié: (2026)
par: Brett, Marshall
Publié: (2026)
Causality $\neq$ Invariance: Function and Concept Vectors in LLMs
par: Opiełka, Gustaw, et autres
Publié: (2026)
par: Opiełka, Gustaw, et autres
Publié: (2026)
Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
par: Frank, Gregory N.
Publié: (2026)
par: Frank, Gregory N.
Publié: (2026)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
par: Hadeliya, Tsimur, et autres
Publié: (2025)
par: Hadeliya, Tsimur, et autres
Publié: (2025)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
par: Cheng, Stephen, et autres
Publié: (2026)
par: Cheng, Stephen, et autres
Publié: (2026)
Why LLMs Cannot Think and How to Fix It
par: Jahrens, Marius, et autres
Publié: (2025)
par: Jahrens, Marius, et autres
Publié: (2025)
Binary Sparse Coding for Interpretability
par: Quirke, Lucia, et autres
Publié: (2025)
par: Quirke, Lucia, et autres
Publié: (2025)
Transcoders Beat Sparse Autoencoders for Interpretability
par: Paulo, Gonçalo, et autres
Publié: (2025)
par: Paulo, Gonçalo, et autres
Publié: (2025)
Training LLMs over Neurally Compressed Text
par: Lester, Brian, et autres
Publié: (2024)
par: Lester, Brian, et autres
Publié: (2024)
Documents similaires
-
Understanding Gradient Descent through the Training Jacobian
par: Belrose, Nora, et autres
Publié: (2024) -
Estimating the Probability of Sampling a Trained Neural Network at Random
par: Scherlis, Adam, et autres
Publié: (2025) -
Does Transformer Interpretability Transfer to RNNs?
par: Paulo, Gonçalo, et autres
Publié: (2024) -
Partially Rewriting a Transformer in Natural Language
par: Paulo, Gonçalo, et autres
Publié: (2025) -
Mechanistic Anomaly Detection for "Quirky" Language Models
par: Johnston, David O., et autres
Publié: (2025)