Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
Fuente:
arXiv
Saved in:
| Main Authors: | Bushnaq, Lucius, Mendel, Jake, Heimersheim, Stefan, Braun, Dan, Goldowsky-Dill, Nicholas, Hänni, Kaarel, Wu, Cindy, Hobbhahn, Marius |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks
by: Bushnaq, Lucius, et al.
Published: (2024)
by: Bushnaq, Lucius, et al.
Published: (2024)
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
by: Braun, Dan, et al.
Published: (2025)
by: Braun, Dan, et al.
Published: (2025)
Detecting Strategic Deception Using Linear Probes
by: Goldowsky-Dill, Nicholas, et al.
Published: (2025)
by: Goldowsky-Dill, Nicholas, et al.
Published: (2025)
Mathematical Models of Computation in Superposition
by: Hänni, Kaarel, et al.
Published: (2024)
by: Hänni, Kaarel, et al.
Published: (2024)
Identifying Sparsely Active Circuits Through Local Loss Landscape Decomposition
by: Chrisman, Brianna, et al.
Published: (2025)
by: Chrisman, Brianna, et al.
Published: (2025)
Stochastic Parameter Decomposition
by: Bushnaq, Lucius, et al.
Published: (2025)
by: Bushnaq, Lucius, et al.
Published: (2025)
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning
by: Braun, Dan, et al.
Published: (2024)
by: Braun, Dan, et al.
Published: (2024)
Open Problems in Mechanistic Interpretability
by: Sharkey, Lee, et al.
Published: (2025)
by: Sharkey, Lee, et al.
Published: (2025)
From Memorization to Reasoning in the Spectrum of Loss Curvature
by: Merullo, Jack, et al.
Published: (2025)
by: Merullo, Jack, et al.
Published: (2025)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
by: Baroni, Luca, et al.
Published: (2025)
by: Baroni, Luca, et al.
Published: (2025)
You can remove GPT2's LayerNorm by fine-tuning
by: Heimersheim, Stefan
Published: (2024)
by: Heimersheim, Stefan
Published: (2024)
Loss Landscape Degeneracy and Stagewise Development in Transformers
by: Hoogland, Jesse, et al.
Published: (2024)
by: Hoogland, Jesse, et al.
Published: (2024)
Cluster-norm for Unsupervised Probing of Knowledge
by: Laurito, Walter, et al.
Published: (2024)
by: Laurito, Walter, et al.
Published: (2024)
How to use and interpret activation patching
by: Heimersheim, Stefan, et al.
Published: (2024)
by: Heimersheim, Stefan, et al.
Published: (2024)
Towards evaluations-based safety cases for AI scheming
by: Balesni, Mikita, et al.
Published: (2024)
by: Balesni, Mikita, et al.
Published: (2024)
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
by: Lee, Daniel J., et al.
Published: (2024)
by: Lee, Daniel J., et al.
Published: (2024)
Large Language Models can Strategically Deceive their Users when Put Under Pressure
by: Scheurer, Jérémy, et al.
Published: (2023)
by: Scheurer, Jérémy, et al.
Published: (2023)
Technical Report: Evaluating Goal Drift in Language Model Agents
by: Arike, Rauno, et al.
Published: (2025)
by: Arike, Rauno, et al.
Published: (2025)
A Mechanistic Account of Attention Sinks in GPT-2: One Circuit, Broader Implications for Mitigation
by: Ran-Milo, Yuval, et al.
Published: (2026)
by: Ran-Milo, Yuval, et al.
Published: (2026)
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
by: Taufeeque, Mohammad, et al.
Published: (2026)
by: Taufeeque, Mohammad, et al.
Published: (2026)
Evolution of SAE Features Across Layers in LLMs
by: Balcells, Daniel, et al.
Published: (2024)
by: Balcells, Daniel, et al.
Published: (2024)
Frontier Models are Capable of In-context Scheming
by: Meinke, Alexander, et al.
Published: (2024)
by: Meinke, Alexander, et al.
Published: (2024)
Exemplar Partitioning for Mechanistic Interpretability
by: Rumbelow, Jessica
Published: (2026)
by: Rumbelow, Jessica
Published: (2026)
From Mechanistic to Compositional Interpretability
by: Gauderis, Ward, et al.
Published: (2026)
by: Gauderis, Ward, et al.
Published: (2026)
Evaluating Synthetic Activations composed of SAE Latents in GPT-2
by: Giglemiani, Giorgi, et al.
Published: (2024)
by: Giglemiani, Giorgi, et al.
Published: (2024)
Landscaper: Understanding Loss Landscapes Through Multi-Dimensional Topological Analysis
by: Chen, Jiaqing, et al.
Published: (2026)
by: Chen, Jiaqing, et al.
Published: (2026)
Sensitivity Analysis On Loss Landscape
by: Faroz, Salman
Published: (2024)
by: Faroz, Salman
Published: (2024)
Mechanistic Interpretability for Neural TSP Solvers
by: Narad, Reuben, et al.
Published: (2025)
by: Narad, Reuben, et al.
Published: (2025)
Mechanistic Interpretability of Reinforcement Learning Agents
by: Trim, Tristan, et al.
Published: (2024)
by: Trim, Tristan, et al.
Published: (2024)
Validating Mechanistic Interpretations: An Axiomatic Approach
by: Palumbo, Nils, et al.
Published: (2024)
by: Palumbo, Nils, et al.
Published: (2024)
Visualizing Critic Match Loss Landscapes for Interpretation of Online Reinforcement Learning Control Algorithms
by: Liu, Jingyi, et al.
Published: (2026)
by: Liu, Jingyi, et al.
Published: (2026)
A Loss Landscape Visualization Framework for Interpreting Reinforcement Learning: An ADHDP Case Study
by: Liu, Jingyi, et al.
Published: (2026)
by: Liu, Jingyi, et al.
Published: (2026)
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
by: Kowal, Matthew, et al.
Published: (2026)
by: Kowal, Matthew, et al.
Published: (2026)
Geospatial Mechanistic Interpretability of Large Language Models
by: De Sabbata, Stef, et al.
Published: (2025)
by: De Sabbata, Stef, et al.
Published: (2025)
Benchmarking Deception Probes via Black-to-White Performance Boosts
by: Parrack, Avi, et al.
Published: (2025)
by: Parrack, Avi, et al.
Published: (2025)
Characterizing stable regions in the residual stream of LLMs
by: Janiak, Jett, et al.
Published: (2024)
by: Janiak, Jett, et al.
Published: (2024)
SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs
by: Fillingham, Sean P., et al.
Published: (2025)
by: Fillingham, Sean P., et al.
Published: (2025)
There is a Singularity in the Loss Landscape
by: Lowell, Mark
Published: (2022)
by: Lowell, Mark
Published: (2022)
Paths and Ambient Spaces in Neural Loss Landscapes
by: Dold, Daniel, et al.
Published: (2025)
by: Dold, Daniel, et al.
Published: (2025)
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
by: Saini, Harshvardhan, et al.
Published: (2026)
by: Saini, Harshvardhan, et al.
Published: (2026)
Similar Items
-
The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks
by: Bushnaq, Lucius, et al.
Published: (2024) -
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
by: Braun, Dan, et al.
Published: (2025) -
Detecting Strategic Deception Using Linear Probes
by: Goldowsky-Dill, Nicholas, et al.
Published: (2025) -
Mathematical Models of Computation in Superposition
by: Hänni, Kaarel, et al.
Published: (2024) -
Identifying Sparsely Active Circuits Through Local Loss Landscape Decomposition
by: Chrisman, Brianna, et al.
Published: (2025)