Identifying Sparsely Active Circuits Through Local Loss Landscape Decomposition
Fuente:
arXiv
Saved in:
| Main Authors: | Chrisman, Brianna, Bushnaq, Lucius, Sharkey, Lee |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stochastic Parameter Decomposition
by: Bushnaq, Lucius, et al.
Published: (2025)
by: Bushnaq, Lucius, et al.
Published: (2025)
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
by: Braun, Dan, et al.
Published: (2025)
by: Braun, Dan, et al.
Published: (2025)
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning
by: Braun, Dan, et al.
Published: (2024)
by: Braun, Dan, et al.
Published: (2024)
From Memorization to Reasoning in the Spectrum of Loss Curvature
by: Merullo, Jack, et al.
Published: (2025)
by: Merullo, Jack, et al.
Published: (2025)
Landscaper: Understanding Loss Landscapes Through Multi-Dimensional Topological Analysis
by: Chen, Jiaqing, et al.
Published: (2026)
by: Chen, Jiaqing, et al.
Published: (2026)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
by: Bushnaq, Lucius, et al.
Published: (2024)
by: Bushnaq, Lucius, et al.
Published: (2024)
Sparse Attention Decomposition Applied to Circuit Tracing
by: Franco, Gabriel, et al.
Published: (2024)
by: Franco, Gabriel, et al.
Published: (2024)
OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition
by: Zhang, Stephen, et al.
Published: (2024)
by: Zhang, Stephen, et al.
Published: (2024)
The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks
by: Bushnaq, Lucius, et al.
Published: (2024)
by: Bushnaq, Lucius, et al.
Published: (2024)
Sparse Autoencoders Do Not Find Canonical Units of Analysis
by: Leask, Patrick, et al.
Published: (2025)
by: Leask, Patrick, et al.
Published: (2025)
There is a Singularity in the Loss Landscape
by: Lowell, Mark
Published: (2022)
by: Lowell, Mark
Published: (2022)
Sparse Probabilistic Graph Circuits
by: Rektoris, Martin, et al.
Published: (2025)
by: Rektoris, Martin, et al.
Published: (2025)
Visualizing Loss Functions as Topological Landscape Profiles
by: Geniesse, Caleb, et al.
Published: (2024)
by: Geniesse, Caleb, et al.
Published: (2024)
Sparse Decomposition of Graph Neural Networks
by: Hu, Yaochen, et al.
Published: (2024)
by: Hu, Yaochen, et al.
Published: (2024)
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
by: Ayonrinde, Kola, et al.
Published: (2024)
by: Ayonrinde, Kola, et al.
Published: (2024)
Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs
by: Ye, Charles, et al.
Published: (2026)
by: Ye, Charles, et al.
Published: (2026)
Model Merging on Loss Landscape: A Geometry Perspective
by: Lu, Juanwu, et al.
Published: (2026)
by: Lu, Juanwu, et al.
Published: (2026)
Evaluating Loss Landscapes from a Topology Perspective
by: Xie, Tiankai, et al.
Published: (2024)
by: Xie, Tiankai, et al.
Published: (2024)
Matrix Sensing with Kernel Optimal Loss: Robustness and Optimization Landscape
by: Song, Xinyuan, et al.
Published: (2025)
by: Song, Xinyuan, et al.
Published: (2025)
Early-Warning Signals of Grokking via Loss-Landscape Geometry
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Hierarchical Sparse Circuit Extraction from Billion-Parameter Language Models through Scalable Attribution Graph Decomposition
by: Uddin, Mohammed Mudassir, et al.
Published: (2026)
by: Uddin, Mohammed Mudassir, et al.
Published: (2026)
A Tale of Two Symmetries: Exploring the Loss Landscape of Equivariant Models
by: Xie, YuQing, et al.
Published: (2025)
by: Xie, YuQing, et al.
Published: (2025)
Practical Bayesian Inference for Speech SNNs: Uncertainty and Loss-Landscape Smoothing
by: Abdennadher, Yesmine, et al.
Published: (2026)
by: Abdennadher, Yesmine, et al.
Published: (2026)
Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Locality Sensitive Sparse Encoding for Learning World Models Online
by: Liu, Zichen, et al.
Published: (2024)
by: Liu, Zichen, et al.
Published: (2024)
Finding the Muses: Identifying Coresets through Loss Trajectories
by: Nagaraj, Manish, et al.
Published: (2025)
by: Nagaraj, Manish, et al.
Published: (2025)
Adaptive Weighted Loss for Sequential Recommendations on Sparse Domains
by: Mittal, Akshay, et al.
Published: (2025)
by: Mittal, Akshay, et al.
Published: (2025)
Loss Landscape Degeneracy and Stagewise Development in Transformers
by: Hoogland, Jesse, et al.
Published: (2024)
by: Hoogland, Jesse, et al.
Published: (2024)
On Discovery of Local Independence over Continuous Variables via Neural Contextual Decomposition
by: Hwang, Inwoo, et al.
Published: (2024)
by: Hwang, Inwoo, et al.
Published: (2024)
Semantic Optimal Transport for Sparse Autoencoder Feature Matching and Circuit Compression
by: Cao, Tue M., et al.
Published: (2026)
by: Cao, Tue M., et al.
Published: (2026)
On the Convergence of Loss and Uncertainty-based Active Learning Algorithms
by: Haimovich, Daniel, et al.
Published: (2023)
by: Haimovich, Daniel, et al.
Published: (2023)
Language Models as Causal Effect Generators
by: Bynum, Lucius E. J., et al.
Published: (2024)
by: Bynum, Lucius E. J., et al.
Published: (2024)
SLaB: Sparse-Lowrank-Binary Decomposition for Efficient Large Language Models
by: Li, Ziwei, et al.
Published: (2026)
by: Li, Ziwei, et al.
Published: (2026)
Fast Adversarial Training against Sparse Attacks Requires Loss Smoothing
by: Zhong, Xuyang, et al.
Published: (2025)
by: Zhong, Xuyang, et al.
Published: (2025)
LADDER: Self-Improving LLMs Through Recursive Problem Decomposition
by: Simonds, Toby, et al.
Published: (2025)
by: Simonds, Toby, et al.
Published: (2025)
A New Paradigm for Counterfactual Reasoning in Fairness and Recourse
by: Bynum, Lucius E. J., et al.
Published: (2024)
by: Bynum, Lucius E. J., et al.
Published: (2024)
Flat Channels to Infinity in Neural Loss Landscapes
by: Martinelli, Flavio, et al.
Published: (2025)
by: Martinelli, Flavio, et al.
Published: (2025)
Adapting Critic Match Loss Landscape Visualization to Off-policy Reinforcement Learning
by: Liu, Jingyi, et al.
Published: (2026)
by: Liu, Jingyi, et al.
Published: (2026)
Loss Barcode: A Topological Measure of Escapability in Loss Landscapes
by: Barannikov, Serguei, et al.
Published: (2020)
by: Barannikov, Serguei, et al.
Published: (2020)
Efficient Automated Circuit Discovery in Transformers using Contextual Decomposition
by: Hsu, Aliyah R., et al.
Published: (2024)
by: Hsu, Aliyah R., et al.
Published: (2024)
Similar Items
-
Stochastic Parameter Decomposition
by: Bushnaq, Lucius, et al.
Published: (2025) -
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
by: Braun, Dan, et al.
Published: (2025) -
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning
by: Braun, Dan, et al.
Published: (2024) -
From Memorization to Reasoning in the Spectrum of Loss Curvature
by: Merullo, Jack, et al.
Published: (2025) -
Landscaper: Understanding Loss Landscapes Through Multi-Dimensional Topological Analysis
by: Chen, Jiaqing, et al.
Published: (2026)