Applying sparse autoencoders to unlearn knowledge in language models
Fuente:
arXiv
Saved in:
| Main Authors: | Farrell, Eoin, Lau, Yeu-Tong, Conmy, Arthur |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling and evaluating sparse autoencoders
by: Gao, Leo, et al.
Published: (2024)
by: Gao, Leo, et al.
Published: (2024)
Scaling sparse feature circuit finding for in-context learning
by: Kharlapenko, Dmitrii, et al.
Published: (2025)
by: Kharlapenko, Dmitrii, et al.
Published: (2025)
An Adversarial Example for Direct Logit Attribution: Memory Management in GELU-4L
by: Janiak, Jett, et al.
Published: (2023)
by: Janiak, Jett, et al.
Published: (2023)
Automatically Finding Reward Model Biases
by: Wang, Atticus, et al.
Published: (2026)
by: Wang, Atticus, et al.
Published: (2026)
Provable unlearning in topic modeling and downstream tasks
by: Wei, Stanley, et al.
Published: (2024)
by: Wei, Stanley, et al.
Published: (2024)
Machine unlearning through fine-grained model parameters perturbation
by: Zuo, Zhiwei, et al.
Published: (2024)
by: Zuo, Zhiwei, et al.
Published: (2024)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
by: Chalnev, Sviatoslav, et al.
Published: (2024)
by: Chalnev, Sviatoslav, et al.
Published: (2024)
Transformer autoencoder with local attention for sparse and irregular time series with application on risk estimation
by: Rodis, Panteleimon
Published: (2026)
by: Rodis, Panteleimon
Published: (2026)
Can sparse autoencoders be used to decompose and interpret steering vectors?
by: Mayne, Harry, et al.
Published: (2024)
by: Mayne, Harry, et al.
Published: (2024)
Steering CLIP's vision transformer with sparse autoencoders
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
Understanding Reasoning in Thinking Language Models via Steering Vectors
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
Base Models Know How to Reason, Thinking Models Learn When
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
Training data attribution in diffusion models via mirrored unlearning and noise-consistent skew
by: Serrà, Joan, et al.
Published: (2026)
by: Serrà, Joan, et al.
Published: (2026)
Thought Anchors: Which LLM Reasoning Steps Matter?
by: Bogdan, Paul C., et al.
Published: (2025)
by: Bogdan, Paul C., et al.
Published: (2025)
On the limitation of evaluating machine unlearning using only a single training seed
by: Lanyon, Jamie, et al.
Published: (2025)
by: Lanyon, Jamie, et al.
Published: (2025)
Variational autoencoder-based neural network model compression
by: Cheng, Liang, et al.
Published: (2024)
by: Cheng, Liang, et al.
Published: (2024)
Multimodal normative modeling in Alzheimers Disease with introspective variational autoencoders
by: Kumar, Sayantan, et al.
Published: (2026)
by: Kumar, Sayantan, et al.
Published: (2026)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
by: Arcuschin, Iván, et al.
Published: (2025)
by: Arcuschin, Iván, et al.
Published: (2025)
How do LLMs Compute Verbal Confidence
by: Kumaran, Dharshan, et al.
Published: (2026)
by: Kumaran, Dharshan, et al.
Published: (2026)
Improving Dictionary Learning with Gated Sparse Autoencoders
by: Rajamanoharan, Senthooran, et al.
Published: (2024)
by: Rajamanoharan, Senthooran, et al.
Published: (2024)
Discriminative protein sequence modelling with Latent Space Diffusion
by: Quinn, Eoin, et al.
Published: (2025)
by: Quinn, Eoin, et al.
Published: (2025)
Physics-integrated generative modeling using attentive planar normalizing flow based variational autoencoder
by: Akhtar, Sheikh Waqas
Published: (2024)
by: Akhtar, Sheikh Waqas
Published: (2024)
Building Production-Ready Probes For Gemini
by: Kramár, János, et al.
Published: (2026)
by: Kramár, János, et al.
Published: (2026)
Beyond designer's knowledge: Generating materials design hypotheses via large language models
by: Liu, Quanliang, et al.
Published: (2024)
by: Liu, Quanliang, et al.
Published: (2024)
Epistemic diversity across language models mitigates knowledge collapse
by: Hodel, Damian, et al.
Published: (2025)
by: Hodel, Damian, et al.
Published: (2025)
Pre-trained knowledge elevates large language models beyond traditional chemical reaction optimizers
by: MacKnight, Robert, et al.
Published: (2025)
by: MacKnight, Robert, et al.
Published: (2025)
Remote sensing framework for geological mapping via stacked autoencoders and clustering
by: Nagar, Sandeep, et al.
Published: (2024)
by: Nagar, Sandeep, et al.
Published: (2024)
Towards efficient deep autoencoders for multivariate time series anomaly detection
by: Pietroń, Marcin, et al.
Published: (2024)
by: Pietroń, Marcin, et al.
Published: (2024)
A practical existence theorem for reduced order models based on convolutional autoencoders
by: Franco, Nicola Rares, et al.
Published: (2024)
by: Franco, Nicola Rares, et al.
Published: (2024)
Insights into a radiology-specialised multimodal large language model with sparse autoencoders
by: Bouzid, Kenza, et al.
Published: (2025)
by: Bouzid, Kenza, et al.
Published: (2025)
Discretization of continuous input spaces in the hippocampal autoencoder
by: Amil, Adrian F., et al.
Published: (2024)
by: Amil, Adrian F., et al.
Published: (2024)
Multi-view autoencoders for Fake News Detection
by: Pereira, Ingryd V. S. T., et al.
Published: (2025)
by: Pereira, Ingryd V. S. T., et al.
Published: (2025)
How predictable is language model benchmark performance?
by: Owen, David
Published: (2024)
by: Owen, David
Published: (2024)
Metabolic cost of information processing in Poisson variational autoencoders
by: Vafaii, Hadi, et al.
Published: (2026)
by: Vafaii, Hadi, et al.
Published: (2026)
Efficiency optimization of large-scale language models based on deep learning in natural language processing tasks
by: Mei, Taiyuan, et al.
Published: (2024)
by: Mei, Taiyuan, et al.
Published: (2024)
Quantifying construct validity in large language model evaluations
by: Kearns, Ryan Othniel
Published: (2026)
by: Kearns, Ryan Othniel
Published: (2026)
Weight-sparse transformers have interpretable circuits
by: Gao, Leo, et al.
Published: (2025)
by: Gao, Leo, et al.
Published: (2025)
Inferring response times of perceptual decisions with Poisson variational autoencoders
by: Johnson, Hayden R., et al.
Published: (2025)
by: Johnson, Hayden R., et al.
Published: (2025)
Well2Flow: Reconstruction of reservoir states from sparse wells using score-based generative models
by: Zeng, Shiqin, et al.
Published: (2025)
by: Zeng, Shiqin, et al.
Published: (2025)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
by: Lieberum, Tom, et al.
Published: (2024)
by: Lieberum, Tom, et al.
Published: (2024)
Similar Items
-
Scaling and evaluating sparse autoencoders
by: Gao, Leo, et al.
Published: (2024) -
Scaling sparse feature circuit finding for in-context learning
by: Kharlapenko, Dmitrii, et al.
Published: (2025) -
An Adversarial Example for Direct Logit Attribution: Memory Management in GELU-4L
by: Janiak, Jett, et al.
Published: (2023) -
Automatically Finding Reward Model Biases
by: Wang, Atticus, et al.
Published: (2026) -
Provable unlearning in topic modeling and downstream tasks
by: Wei, Stanley, et al.
Published: (2024)