How to use and interpret activation patching
Fuente:
arXiv
Guardado en:
| Autores principales: | Heimersheim, Stefan, Nanda, Neel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
You can remove GPT2's LayerNorm by fine-tuning
por: Heimersheim, Stefan
Publicado: (2024)
por: Heimersheim, Stefan
Publicado: (2024)
Towards eliciting latent knowledge from LLMs with mechanistic interpretability
por: Cywiński, Bartosz, et al.
Publicado: (2025)
por: Cywiński, Bartosz, et al.
Publicado: (2025)
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
por: Lee, Daniel J., et al.
Publicado: (2024)
por: Lee, Daniel J., et al.
Publicado: (2024)
Explorations of Self-Repair in Language Models
por: Rushing, Cody, et al.
Publicado: (2024)
por: Rushing, Cody, et al.
Publicado: (2024)
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
por: Zhang, Fred, et al.
Publicado: (2023)
por: Zhang, Fred, et al.
Publicado: (2023)
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
por: Makelov, Aleksandar, et al.
Publicado: (2024)
por: Makelov, Aleksandar, et al.
Publicado: (2024)
Difficulties with Evaluating a Deception Detector for AIs
por: Smith, Lewis, et al.
Publicado: (2025)
por: Smith, Lewis, et al.
Publicado: (2025)
Base Models Know How to Reason, Thinking Models Learn When
por: Venhoff, Constantin, et al.
Publicado: (2025)
por: Venhoff, Constantin, et al.
Publicado: (2025)
Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models
por: Leask, Patrick, et al.
Publicado: (2025)
por: Leask, Patrick, et al.
Publicado: (2025)
BatchTopK Sparse Autoencoders
por: Bussmann, Bart, et al.
Publicado: (2024)
por: Bussmann, Bart, et al.
Publicado: (2024)
Transcoders Find Interpretable LLM Feature Circuits
por: Dunefsky, Jacob, et al.
Publicado: (2024)
por: Dunefsky, Jacob, et al.
Publicado: (2024)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
por: Chughtai, Bilal, et al.
Publicado: (2024)
por: Chughtai, Bilal, et al.
Publicado: (2024)
Detecting Strategic Deception Using Linear Probes
por: Goldowsky-Dill, Nicholas, et al.
Publicado: (2025)
por: Goldowsky-Dill, Nicholas, et al.
Publicado: (2025)
How Visual Representations Map to Language Feature Space in Multimodal LLMs
por: Venhoff, Constantin, et al.
Publicado: (2025)
por: Venhoff, Constantin, et al.
Publicado: (2025)
Reasoning-Finetuning Repurposes Latent Representations in Base Models
por: Ward, Jake, et al.
Publicado: (2025)
por: Ward, Jake, et al.
Publicado: (2025)
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
por: Taufeeque, Mohammad, et al.
Publicado: (2026)
por: Taufeeque, Mohammad, et al.
Publicado: (2026)
Evolution of SAE Features Across Layers in LLMs
por: Balcells, Daniel, et al.
Publicado: (2024)
por: Balcells, Daniel, et al.
Publicado: (2024)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
por: Baroni, Luca, et al.
Publicado: (2025)
por: Baroni, Luca, et al.
Publicado: (2025)
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
por: Braun, Dan, et al.
Publicado: (2025)
por: Braun, Dan, et al.
Publicado: (2025)
RelP: Faithful and Efficient Circuit Discovery in Language Models via Relevance Patching
por: Jafari, Farnoush Rezaei, et al.
Publicado: (2025)
por: Jafari, Farnoush Rezaei, et al.
Publicado: (2025)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
por: Karvonen, Adam, et al.
Publicado: (2024)
por: Karvonen, Adam, et al.
Publicado: (2024)
AtP*: An efficient and scalable method for localizing LLM behaviour to components
por: Kramár, János, et al.
Publicado: (2024)
por: Kramár, János, et al.
Publicado: (2024)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
por: Bussmann, Bart, et al.
Publicado: (2025)
por: Bussmann, Bart, et al.
Publicado: (2025)
Convergent Linear Representations of Emergent Misalignment
por: Soligo, Anna, et al.
Publicado: (2025)
por: Soligo, Anna, et al.
Publicado: (2025)
Evaluating Synthetic Activations composed of SAE Latents in GPT-2
por: Giglemiani, Giorgi, et al.
Publicado: (2024)
por: Giglemiani, Giorgi, et al.
Publicado: (2024)
Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval
por: Venhoff, Constantin, et al.
Publicado: (2025)
por: Venhoff, Constantin, et al.
Publicado: (2025)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
por: Ferrando, Javier, et al.
Publicado: (2024)
por: Ferrando, Javier, et al.
Publicado: (2024)
Interpreting Attention Layer Outputs with Sparse Autoencoders
por: Kissane, Connor, et al.
Publicado: (2024)
por: Kissane, Connor, et al.
Publicado: (2024)
Benchmarking Deception Probes via Black-to-White Performance Boosts
por: Parrack, Avi, et al.
Publicado: (2025)
por: Parrack, Avi, et al.
Publicado: (2025)
Characterizing stable regions in the residual stream of LLMs
por: Janiak, Jett, et al.
Publicado: (2024)
por: Janiak, Jett, et al.
Publicado: (2024)
SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs
por: Fillingham, Sean P., et al.
Publicado: (2025)
por: Fillingham, Sean P., et al.
Publicado: (2025)
Model Organisms for Emergent Misalignment
por: Turner, Edward, et al.
Publicado: (2025)
por: Turner, Edward, et al.
Publicado: (2025)
Understanding Reasoning in Thinking Language Models via Steering Vectors
por: Venhoff, Constantin, et al.
Publicado: (2025)
por: Venhoff, Constantin, et al.
Publicado: (2025)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
por: Kantamneni, Subhash, et al.
Publicado: (2025)
por: Kantamneni, Subhash, et al.
Publicado: (2025)
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
por: Jiang, Nick, et al.
Publicado: (2025)
por: Jiang, Nick, et al.
Publicado: (2025)
Simple LLM Baselines are Competitive for Model Diffing
por: Kempf, Elias, et al.
Publicado: (2026)
por: Kempf, Elias, et al.
Publicado: (2026)
Thought Branches: Interpreting LLM Reasoning Requires Resampling
por: Macar, Uzay, et al.
Publicado: (2025)
por: Macar, Uzay, et al.
Publicado: (2025)
Thought Anchors: Which LLM Reasoning Steps Matter?
por: Bogdan, Paul C., et al.
Publicado: (2025)
por: Bogdan, Paul C., et al.
Publicado: (2025)
Simple Mechanistic Explanations for Out-Of-Context Reasoning
por: Wang, Atticus, et al.
Publicado: (2025)
por: Wang, Atticus, et al.
Publicado: (2025)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
por: Maar, Jim, et al.
Publicado: (2026)
por: Maar, Jim, et al.
Publicado: (2026)
Ejemplares similares
-
You can remove GPT2's LayerNorm by fine-tuning
por: Heimersheim, Stefan
Publicado: (2024) -
Towards eliciting latent knowledge from LLMs with mechanistic interpretability
por: Cywiński, Bartosz, et al.
Publicado: (2025) -
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
por: Lee, Daniel J., et al.
Publicado: (2024) -
Explorations of Self-Repair in Language Models
por: Rushing, Cody, et al.
Publicado: (2024) -
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
por: Zhang, Fred, et al.
Publicado: (2023)