Residual Stream Analysis with Multi-Layer SAEs
Fuente:
arXiv
Saved in:
| Main Authors: | Lawson, Tim, Farnik, Lucy, Houghton, Conor, Aitchison, Laurence |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
by: Farnik, Lucy, et al.
Published: (2025)
by: Farnik, Lucy, et al.
Published: (2025)
Learning to Skip the Middle Layers of Transformers
by: Lawson, Tim, et al.
Published: (2025)
by: Lawson, Tim, et al.
Published: (2025)
Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
by: Heap, Thomas, et al.
Published: (2025)
by: Heap, Thomas, et al.
Published: (2025)
Scale-invariant Attention
by: Anson, Ben, et al.
Published: (2025)
by: Anson, Ben, et al.
Published: (2025)
Image, Word and Thought: A More Challenging Language Task for the Iterated Learning Model
by: Lee, Hyoyeon, et al.
Published: (2026)
by: Lee, Hyoyeon, et al.
Published: (2026)
Can SAEs reveal and mitigate racial biases of LLMs in healthcare?
by: Ahsan, Hiba, et al.
Published: (2025)
by: Ahsan, Hiba, et al.
Published: (2025)
Resa: Transparent Reasoning Models via SAEs
by: Wang, Shangshang, et al.
Published: (2025)
by: Wang, Shangshang, et al.
Published: (2025)
Residual Matrix Transformers: Scaling the Size of the Residual Stream
by: Mak, Brian, et al.
Published: (2025)
by: Mak, Brian, et al.
Published: (2025)
SAEs Are Good for Steering -- If You Select the Right Features
by: Arad, Dana, et al.
Published: (2025)
by: Arad, Dana, et al.
Published: (2025)
Teach Old SAEs New Domain Tricks with Boosting
by: Koriagin, Nikita, et al.
Published: (2025)
by: Koriagin, Nikita, et al.
Published: (2025)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
by: Song, Xiangchen, et al.
Published: (2025)
by: Song, Xiangchen, et al.
Published: (2025)
Questionable practices in machine learning
by: Leech, Gavin, et al.
Published: (2024)
by: Leech, Gavin, et al.
Published: (2024)
Residual Stream Duality in Modern Transformer Architectures
by: Zhang, Yifan
Published: (2026)
by: Zhang, Yifan
Published: (2026)
Why you don't overfit, and don't need Bayes if you only train for one epoch
by: Aitchison, Laurence
Published: (2024)
by: Aitchison, Laurence
Published: (2024)
SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection
by: Joshi, Maithili, et al.
Published: (2025)
by: Joshi, Maithili, et al.
Published: (2025)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
by: Su, Guinan, et al.
Published: (2026)
by: Su, Guinan, et al.
Published: (2026)
ResidualTransformer: Residual Low-Rank Learning with Weight-Sharing for Transformer Layers
by: Wang, Yiming, et al.
Published: (2023)
by: Wang, Yiming, et al.
Published: (2023)
Modeling language contact with the Iterated Learning Model
by: Bullock, Seth, et al.
Published: (2024)
by: Bullock, Seth, et al.
Published: (2024)
Multi-Layer Attention is the Amplifier of Demonstration Effectiveness
by: Wang, Dingzirui, et al.
Published: (2025)
by: Wang, Dingzirui, et al.
Published: (2025)
Residual Speech Embeddings for Tone Classification: Removing Linguistic Content to Enhance Paralinguistic Analysis
by: Ahbabi, Hamdan Al, et al.
Published: (2025)
by: Ahbabi, Hamdan Al, et al.
Published: (2025)
Better Prompt Compression Without Multi-Layer Perceptrons
by: Honig, Edouardo, et al.
Published: (2025)
by: Honig, Edouardo, et al.
Published: (2025)
Layer by Layer: Uncovering Where Multi-Task Learning Happens in Instruction-Tuned Large Language Models
by: Zhao, Zheng, et al.
Published: (2024)
by: Zhao, Zheng, et al.
Published: (2024)
Latent Phase-Shift Rollback: Inference-Time Error Correction via Residual Stream Monitoring and KV-Cache Steering
by: Gupta, Manan, et al.
Published: (2026)
by: Gupta, Manan, et al.
Published: (2026)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
by: Musat, Tiberiu
Published: (2024)
by: Musat, Tiberiu
Published: (2024)
Batch size invariant Adam
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Controlling changes to attention logits
by: Anson, Ben, et al.
Published: (2025)
by: Anson, Ben, et al.
Published: (2025)
Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams
by: Llorente-Saguer, Isaac
Published: (2026)
by: Llorente-Saguer, Isaac
Published: (2026)
Investigating the Timescales of Language Processing with EEG and Language Models
by: Turco, Davide, et al.
Published: (2024)
by: Turco, Davide, et al.
Published: (2024)
An Analysis of Embedding Layers and Similarity Scores using Siamese Neural Networks
by: Bingi, Yash, et al.
Published: (2023)
by: Bingi, Yash, et al.
Published: (2023)
Residual-Mass Accounting for Partial-KV Decoding
by: Hoshi, Yasuto, et al.
Published: (2026)
by: Hoshi, Yasuto, et al.
Published: (2026)
Stacking Small Language Models for Generalizability
by: Liang, Laurence
Published: (2024)
by: Liang, Laurence
Published: (2024)
Advancing Multi-Step Mathematical Reasoning in Large Language Models through Multi-Layered Self-Reflection with Auto-Prompting
by: Loureiro, André de Souza, et al.
Published: (2025)
by: Loureiro, André de Souza, et al.
Published: (2025)
SequenceLayers: Sequence Processing and Streaming Neural Networks Made Easy
by: Skerry-Ryan, RJ, et al.
Published: (2025)
by: Skerry-Ryan, RJ, et al.
Published: (2025)
Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations
by: Blumenberg, Patrick, et al.
Published: (2025)
by: Blumenberg, Patrick, et al.
Published: (2025)
An Ensemble Classification Approach in A Multi-Layered Large Language Model Framework for Disease Prediction
by: Hamdi, Ali, et al.
Published: (2025)
by: Hamdi, Ali, et al.
Published: (2025)
LayerBoost: Layer-Aware Attention Reduction for Efficient LLMs
by: Souibgui, Mohamed Ali, et al.
Published: (2026)
by: Souibgui, Mohamed Ali, et al.
Published: (2026)
How to set AdamW's weight decay as you scale model and dataset size
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Online Cascade Learning for Efficient Inference over Streams
by: Nie, Lunyiu, et al.
Published: (2024)
by: Nie, Lunyiu, et al.
Published: (2024)
Investigating the Synergistic Effects of Dropout and Residual Connections on Language Model Training
by: Li, Qingyang, et al.
Published: (2024)
by: Li, Qingyang, et al.
Published: (2024)
Similar Items
-
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
by: Farnik, Lucy, et al.
Published: (2025) -
Learning to Skip the Middle Layers of Transformers
by: Lawson, Tim, et al.
Published: (2025) -
Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
by: Heap, Thomas, et al.
Published: (2025) -
Scale-invariant Attention
by: Anson, Ben, et al.
Published: (2025) -
Image, Word and Thought: A More Challenging Language Task for the Iterated Learning Model
by: Lee, Hyoyeon, et al.
Published: (2026)