Slowing Learning by Erasing Simple Features
Fuente:
arXiv
Saved in:
| Main Authors: | Quirke, Lucia, Belrose, Nora |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Binary Sparse Coding for Interpretability
by: Quirke, Lucia, et al.
Published: (2025)
by: Quirke, Lucia, et al.
Published: (2025)
Neural Networks Learn Statistics of Increasing Complexity
by: Belrose, Nora, et al.
Published: (2024)
by: Belrose, Nora, et al.
Published: (2024)
Sparse Autoencoders Trained on the Same Data Learn Different Features
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Estimating the Probability of Sampling a Trained Neural Network at Random
by: Scherlis, Adam, et al.
Published: (2025)
by: Scherlis, Adam, et al.
Published: (2025)
Evaluating SAE interpretability without explanations
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Converting MLPs into Polynomials in Closed Form
by: Belrose, Nora, et al.
Published: (2025)
by: Belrose, Nora, et al.
Published: (2025)
Balancing Label Quantity and Quality for Scalable Elicitation
by: Mallen, Alex, et al.
Published: (2024)
by: Mallen, Alex, et al.
Published: (2024)
Understanding Gradient Descent through the Training Jacobian
by: Belrose, Nora, et al.
Published: (2024)
by: Belrose, Nora, et al.
Published: (2024)
Partially Rewriting a Transformer in Natural Language
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Examining Two Hop Reasoning Through Information Content Scaling
by: Johnston, David, et al.
Published: (2025)
by: Johnston, David, et al.
Published: (2025)
Automatically Interpreting Millions of Features in Large Language Models
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
Transcoders Beat Sparse Autoencoders for Interpretability
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Refusal in LLMs is an Affine Function
by: Marshall, Thomas, et al.
Published: (2024)
by: Marshall, Thomas, et al.
Published: (2024)
Does Transformer Interpretability Transfer to RNNs?
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
Mechanistic Anomaly Detection for "Quirky" Language Models
by: Johnston, David O., et al.
Published: (2025)
by: Johnston, David O., et al.
Published: (2025)
Eliciting Latent Knowledge from Quirky Language Models
by: Mallen, Alex, et al.
Published: (2023)
by: Mallen, Alex, et al.
Published: (2023)
Understanding Addition in Transformers
by: Quirke, Philip, et al.
Published: (2023)
by: Quirke, Philip, et al.
Published: (2023)
Learning Successor Features the Simple Way
by: Chua, Raymond, et al.
Published: (2024)
by: Chua, Raymond, et al.
Published: (2024)
Understanding Addition and Subtraction in Transformers
by: Quirke, Philip, et al.
Published: (2024)
by: Quirke, Philip, et al.
Published: (2024)
Slow Feature Analysis as Variational Inference Objective
by: Schüler, Merlin, et al.
Published: (2025)
by: Schüler, Merlin, et al.
Published: (2025)
Erasing the Bias: Fine-Tuning Foundation Models for Semi-Supervised Learning
by: Gan, Kai, et al.
Published: (2024)
by: Gan, Kai, et al.
Published: (2024)
Balancing Plasticity and Stability with Fast and Slow Successor Features
by: Chua, Raymond, et al.
Published: (2026)
by: Chua, Raymond, et al.
Published: (2026)
LEACE: Perfect linear concept erasure in closed form
by: Belrose, Nora, et al.
Published: (2023)
by: Belrose, Nora, et al.
Published: (2023)
Eliciting Latent Predictions from Transformers with the Tuned Lens
by: Belrose, Nora, et al.
Published: (2023)
by: Belrose, Nora, et al.
Published: (2023)
What is the relation between Slow Feature Analysis and the Successor Representation?
by: Seabrook, Eddie, et al.
Published: (2024)
by: Seabrook, Eddie, et al.
Published: (2024)
Slow Feature Analysis on Markov Chains from Goal-Directed Behavior
by: Schüler, Merlin, et al.
Published: (2025)
by: Schüler, Merlin, et al.
Published: (2025)
Erase or Hide? Suppressing Spurious Unlearning Neurons for Robust Unlearning
by: Yang, Nakyeong, et al.
Published: (2025)
by: Yang, Nakyeong, et al.
Published: (2025)
Erasing Conceptual Knowledge from Language Models
by: Gandikota, Rohit, et al.
Published: (2024)
by: Gandikota, Rohit, et al.
Published: (2024)
A Simple Data Augmentation for Feature Distribution Skewed Federated Learning
by: Yan, Yunlu, et al.
Published: (2023)
by: Yan, Yunlu, et al.
Published: (2023)
When Are Concepts Erased From Diffusion Models?
by: Lu, Kevin, et al.
Published: (2025)
by: Lu, Kevin, et al.
Published: (2025)
Erase at the Core: Representation Unlearning for Machine Unlearning
by: Lee, Jaewon, et al.
Published: (2026)
by: Lee, Jaewon, et al.
Published: (2026)
Redirection for Erasing Memory (REM): Towards a universal unlearning method for corrupted data
by: Schoepf, Stefan, et al.
Published: (2025)
by: Schoepf, Stefan, et al.
Published: (2025)
Flip Learning: Weakly Supervised Erase to Segment Nodules in Breast Ultrasound
by: Huang, Yuhao, et al.
Published: (2025)
by: Huang, Yuhao, et al.
Published: (2025)
Dichotomy of Feature Learning and Unlearning: Fast-Slow Analysis on Neural Networks with Stochastic Gradient Descent
by: Imai, Shota, et al.
Published: (2026)
by: Imai, Shota, et al.
Published: (2026)
EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven Alignment
by: Kusumba, Abhiram, et al.
Published: (2025)
by: Kusumba, Abhiram, et al.
Published: (2025)
Erasing Without Remembering: Implicit Knowledge Forgetting in Large Language Models
by: Wang, Huazheng, et al.
Published: (2025)
by: Wang, Huazheng, et al.
Published: (2025)
Side Effects of Erasing Concepts from Diffusion Models
by: Saha, Shaswati, et al.
Published: (2025)
by: Saha, Shaswati, et al.
Published: (2025)
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
by: Braun, Tobias, et al.
Published: (2025)
by: Braun, Tobias, et al.
Published: (2025)
Erasing Undesirable Concepts in Diffusion Models with Adversarial Preservation
by: Bui, Anh, et al.
Published: (2024)
by: Bui, Anh, et al.
Published: (2024)
Simulating, Fast and Slow: Learning Policies for Black-Box Optimization
by: Massoli, Fabio Valerio, et al.
Published: (2024)
by: Massoli, Fabio Valerio, et al.
Published: (2024)
Similar Items
-
Binary Sparse Coding for Interpretability
by: Quirke, Lucia, et al.
Published: (2025) -
Neural Networks Learn Statistics of Increasing Complexity
by: Belrose, Nora, et al.
Published: (2024) -
Sparse Autoencoders Trained on the Same Data Learn Different Features
by: Paulo, Gonçalo, et al.
Published: (2025) -
Estimating the Probability of Sampling a Trained Neural Network at Random
by: Scherlis, Adam, et al.
Published: (2025) -
Evaluating SAE interpretability without explanations
by: Paulo, Gonçalo, et al.
Published: (2025)