Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Braun, Dan, Bushnaq, Lucius, Heimersheim, Stefan, Mendel, Jake, Sharkey, Lee |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stochastic Parameter Decomposition
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2025)
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2025)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
Identifying Sparsely Active Circuits Through Local Loss Landscape Decomposition
von: Chrisman, Brianna, et al.
Veröffentlicht: (2025)
von: Chrisman, Brianna, et al.
Veröffentlicht: (2025)
The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
Open Problems in Mechanistic Interpretability
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
From Memorization to Reasoning in the Spectrum of Loss Curvature
von: Merullo, Jack, et al.
Veröffentlicht: (2025)
von: Merullo, Jack, et al.
Veröffentlicht: (2025)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
von: Baroni, Luca, et al.
Veröffentlicht: (2025)
von: Baroni, Luca, et al.
Veröffentlicht: (2025)
You can remove GPT2's LayerNorm by fine-tuning
von: Heimersheim, Stefan
Veröffentlicht: (2024)
von: Heimersheim, Stefan
Veröffentlicht: (2024)
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
von: Lee, Daniel J., et al.
Veröffentlicht: (2024)
von: Lee, Daniel J., et al.
Veröffentlicht: (2024)
How to use and interpret activation patching
von: Heimersheim, Stefan, et al.
Veröffentlicht: (2024)
von: Heimersheim, Stefan, et al.
Veröffentlicht: (2024)
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning
von: Braun, Dan, et al.
Veröffentlicht: (2024)
von: Braun, Dan, et al.
Veröffentlicht: (2024)
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
von: Kowal, Matthew, et al.
Veröffentlicht: (2026)
von: Kowal, Matthew, et al.
Veröffentlicht: (2026)
A Mechanistic Account of Attention Sinks in GPT-2: One Circuit, Broader Implications for Mitigation
von: Ran-Milo, Yuval, et al.
Veröffentlicht: (2026)
von: Ran-Milo, Yuval, et al.
Veröffentlicht: (2026)
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2024)
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2024)
Adaptive MSD-Splitting: Enhancing C4.5 and Random Forests for Skewed Continuous Attributes
von: Lee, Jake
Veröffentlicht: (2026)
von: Lee, Jake
Veröffentlicht: (2026)
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
von: Joseph, Sonia, et al.
Veröffentlicht: (2025)
von: Joseph, Sonia, et al.
Veröffentlicht: (2025)
Tackling Noisy Labels with Network Parameter Additive Decomposition
von: Wang, Jingyi, et al.
Veröffentlicht: (2024)
von: Wang, Jingyi, et al.
Veröffentlicht: (2024)
Parameter Space Analysis through Guided Visual Interpolations
von: Kantz, Benedikt, et al.
Veröffentlicht: (2025)
von: Kantz, Benedikt, et al.
Veröffentlicht: (2025)
LightSAM: Parameter-Agnostic Sharpness-Aware Minimization
von: Cheng, Yifei, et al.
Veröffentlicht: (2025)
von: Cheng, Yifei, et al.
Veröffentlicht: (2025)
Mathematical Models of Computation in Superposition
von: Hänni, Kaarel, et al.
Veröffentlicht: (2024)
von: Hänni, Kaarel, et al.
Veröffentlicht: (2024)
Parameter-Free Algorithms for Performative Regret Minimization under Decision-Dependent Distributions
von: Park, Sungwoo, et al.
Veröffentlicht: (2024)
von: Park, Sungwoo, et al.
Veröffentlicht: (2024)
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
von: Zhang, Shichang, et al.
Veröffentlicht: (2025)
von: Zhang, Shichang, et al.
Veröffentlicht: (2025)
Detecting Strategic Deception Using Linear Probes
von: Goldowsky-Dill, Nicholas, et al.
Veröffentlicht: (2025)
von: Goldowsky-Dill, Nicholas, et al.
Veröffentlicht: (2025)
State-offset Tuning: State-based Parameter-Efficient Fine-Tuning for State Space Models
von: Kang, Wonjun, et al.
Veröffentlicht: (2025)
von: Kang, Wonjun, et al.
Veröffentlicht: (2025)
Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
von: Chen, Jianhui, et al.
Veröffentlicht: (2026)
von: Chen, Jianhui, et al.
Veröffentlicht: (2026)
Expand Neurons, Not Parameters
von: Kong, Linghao, et al.
Veröffentlicht: (2025)
von: Kong, Linghao, et al.
Veröffentlicht: (2025)
Hierarchical Sparse Circuit Extraction from Billion-Parameter Language Models through Scalable Attribution Graph Decomposition
von: Uddin, Mohammed Mudassir, et al.
Veröffentlicht: (2026)
von: Uddin, Mohammed Mudassir, et al.
Veröffentlicht: (2026)
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2026)
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2026)
The Horcrux: Mechanistically Interpretable Task Decomposition for Detecting and Mitigating Reward Hacking in Embodied AI Systems
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
Learning to Weight Parameters for Training Data Attribution
von: Li, Shuangqi, et al.
Veröffentlicht: (2025)
von: Li, Shuangqi, et al.
Veröffentlicht: (2025)
Parameter-Efficient Fine-Tuning of State Space Models
von: Galim, Kevin, et al.
Veröffentlicht: (2024)
von: Galim, Kevin, et al.
Veröffentlicht: (2024)
Learning Fine-grained Parameter Sharing via Sparse Tensor Decomposition
von: Üyük, Cem, et al.
Veröffentlicht: (2024)
von: Üyük, Cem, et al.
Veröffentlicht: (2024)
Minimizing False-Positive Attributions in Explanations of Non-Linear Models
von: Gjølbye, Anders, et al.
Veröffentlicht: (2025)
von: Gjølbye, Anders, et al.
Veröffentlicht: (2025)
Evolution of SAE Features Across Layers in LLMs
von: Balcells, Daniel, et al.
Veröffentlicht: (2024)
von: Balcells, Daniel, et al.
Veröffentlicht: (2024)
Symmetry in Neural Network Parameter Spaces
von: Zhao, Bo, et al.
Veröffentlicht: (2025)
von: Zhao, Bo, et al.
Veröffentlicht: (2025)
Catalyst: a Novel Regularizer for Structured Pruning with Auxiliary Extension of Parameter Space
von: Jung, Jaeheun, et al.
Veröffentlicht: (2025)
von: Jung, Jaeheun, et al.
Veröffentlicht: (2025)
Federated Majorize-Minimization: Beyond Parameter Aggregation
von: Dieuleveut, Aymeric, et al.
Veröffentlicht: (2025)
von: Dieuleveut, Aymeric, et al.
Veröffentlicht: (2025)
On Catastrophic Forgetting in Low-Rank Decomposition-Based Parameter-Efficient Fine-Tuning
von: Ahmad, Muhammad, et al.
Veröffentlicht: (2026)
von: Ahmad, Muhammad, et al.
Veröffentlicht: (2026)
Exemplar Partitioning for Mechanistic Interpretability
von: Rumbelow, Jessica
Veröffentlicht: (2026)
von: Rumbelow, Jessica
Veröffentlicht: (2026)
From Mechanistic to Compositional Interpretability
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Stochastic Parameter Decomposition
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2025) -
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024) -
Identifying Sparsely Active Circuits Through Local Loss Landscape Decomposition
von: Chrisman, Brianna, et al.
Veröffentlicht: (2025) -
The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024) -
Open Problems in Mechanistic Interpretability
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)