Bilinear MLPs enable weight-based mechanistic interpretability
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pearce, Michael T., Dooms, Thomas, Rigg, Alice, Oramas, Jose M., Sharkey, Lee |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Weight-based Decomposition: A Case for Bilinear MLPs
von: Pearce, Michael T., et al.
Veröffentlicht: (2024)
von: Pearce, Michael T., et al.
Veröffentlicht: (2024)
Bilinear autoencoders find interpretable manifolds
von: Dooms, Thomas, et al.
Veröffentlicht: (2026)
von: Dooms, Thomas, et al.
Veröffentlicht: (2026)
Converting MLPs into Polynomials in Closed Form
von: Belrose, Nora, et al.
Veröffentlicht: (2025)
von: Belrose, Nora, et al.
Veröffentlicht: (2025)
Finding Manifolds With Bilinear Autoencoders
von: Dooms, Thomas, et al.
Veröffentlicht: (2025)
von: Dooms, Thomas, et al.
Veröffentlicht: (2025)
Bilinear Convolution Decomposition for Causal RL Interpretability
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2024)
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2024)
Compositionality Unlocks Deep Interpretable Models
von: Dooms, Thomas, et al.
Veröffentlicht: (2025)
von: Dooms, Thomas, et al.
Veröffentlicht: (2025)
Structural Disentanglement in Bilinear MLPs via Architectural Inductive Bias
von: Nema, Ojasva, et al.
Veröffentlicht: (2026)
von: Nema, Ojasva, et al.
Veröffentlicht: (2026)
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2024)
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2024)
Tokenized SAEs: Disentangling SAE Reconstructions
von: Dooms, Thomas, et al.
Veröffentlicht: (2025)
von: Dooms, Thomas, et al.
Veröffentlicht: (2025)
A mechanistically interpretable neural network for regulatory genomics
von: Tseng, Alex M., et al.
Veröffentlicht: (2024)
von: Tseng, Alex M., et al.
Veröffentlicht: (2024)
An introduction to graphical tensor notation for mechanistic interpretability
von: Taylor, Jordan K.
Veröffentlicht: (2024)
von: Taylor, Jordan K.
Veröffentlicht: (2024)
Distribution-Aware Feature Selection for SAEs
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
Fractals made Practical: Denoising Diffusion as Partitioned Iterated Function Systems
von: Dooms, Ann
Veröffentlicht: (2026)
von: Dooms, Ann
Veröffentlicht: (2026)
Towards eliciting latent knowledge from LLMs with mechanistic interpretability
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
Opening the AI black box: program synthesis via mechanistic interpretability
von: Michaud, Eric J., et al.
Veröffentlicht: (2024)
von: Michaud, Eric J., et al.
Veröffentlicht: (2024)
Bilinear representation mitigates reversal curse and enables consistent model editing
von: Kim, Dong-Kyum, et al.
Veröffentlicht: (2025)
von: Kim, Dong-Kyum, et al.
Veröffentlicht: (2025)
Towards the Characterization of Representations Learned via Capsule-based Network Architectures
von: Tawalbeh, Saja, et al.
Veröffentlicht: (2023)
von: Tawalbeh, Saja, et al.
Veröffentlicht: (2023)
Efficient Post-Hoc Uncertainty Calibration via Variance-Based Smoothing
von: Denoodt, Fabian, et al.
Veröffentlicht: (2025)
von: Denoodt, Fabian, et al.
Veröffentlicht: (2025)
Interpolated-MLPs: Controllable Inductive Bias
von: Wu, Sean, et al.
Veröffentlicht: (2024)
von: Wu, Sean, et al.
Veröffentlicht: (2024)
MLPs at the EOC: Spectrum of the NTK
von: Terjék, Dávid, et al.
Veröffentlicht: (2025)
von: Terjék, Dávid, et al.
Veröffentlicht: (2025)
MLPs at the EOC: Concentration of the NTK
von: Terjék, Dávid, et al.
Veröffentlicht: (2025)
von: Terjék, Dávid, et al.
Veröffentlicht: (2025)
Deep Model Interpretation with Limited Data : A Coreset-based Approach
von: Behzadi-Khormouji, Hamed, et al.
Veröffentlicht: (2024)
von: Behzadi-Khormouji, Hamed, et al.
Veröffentlicht: (2024)
Benchmarking Optimizers for MLPs in Tabular Deep Learning
von: Gorishniy, Yury, et al.
Veröffentlicht: (2026)
von: Gorishniy, Yury, et al.
Veröffentlicht: (2026)
From Mechanistic to Compositional Interpretability
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
Embedding Hardware Approximations in Discrete Genetic-based Training for Printed MLPs
von: Afentaki, Florentia, et al.
Veröffentlicht: (2024)
von: Afentaki, Florentia, et al.
Veröffentlicht: (2024)
Detecting and Characterizing Planning in Language Models
von: Nainani, Jatin, et al.
Veröffentlicht: (2025)
von: Nainani, Jatin, et al.
Veröffentlicht: (2025)
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
von: Braun, Dan, et al.
Veröffentlicht: (2025)
von: Braun, Dan, et al.
Veröffentlicht: (2025)
Locating acts of mechanistic reasoning in student team conversations with mechanistic machine learning
von: Gili, Kaitlin, et al.
Veröffentlicht: (2026)
von: Gili, Kaitlin, et al.
Veröffentlicht: (2026)
Spontaneous Kolmogorov-Arnold Geometry in Shallow MLPs
von: Freedman, Michael H., et al.
Veröffentlicht: (2025)
von: Freedman, Michael H., et al.
Veröffentlicht: (2025)
Hyperparameter Tuning MLPs for Probabilistic Time Series Forecasting
von: Madhusudhanan, Kiran, et al.
Veröffentlicht: (2024)
von: Madhusudhanan, Kiran, et al.
Veröffentlicht: (2024)
Mimetic Initialization of MLPs
von: Trockman, Asher, et al.
Veröffentlicht: (2026)
von: Trockman, Asher, et al.
Veröffentlicht: (2026)
Identifying Sparsely Active Circuits Through Local Loss Landscape Decomposition
von: Chrisman, Brianna, et al.
Veröffentlicht: (2025)
von: Chrisman, Brianna, et al.
Veröffentlicht: (2025)
Stochastic Parameter Decomposition
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2025)
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2025)
Sparse Autoencoders Do Not Find Canonical Units of Analysis
von: Leask, Patrick, et al.
Veröffentlicht: (2025)
von: Leask, Patrick, et al.
Veröffentlicht: (2025)
Constructing Efficient Fact-Storing MLPs for Transformers
von: Dugan, Owen, et al.
Veröffentlicht: (2025)
von: Dugan, Owen, et al.
Veröffentlicht: (2025)
MLPs Learn In-Context on Regression and Classification Tasks
von: Tong, William L., et al.
Veröffentlicht: (2024)
von: Tong, William L., et al.
Veröffentlicht: (2024)
Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPs
von: Chowdhury, Tanya, et al.
Veröffentlicht: (2025)
von: Chowdhury, Tanya, et al.
Veröffentlicht: (2025)
Why Attention Fails: The Degeneration of Transformers into MLPs in Time Series Forecasting
von: Liang, Zida, et al.
Veröffentlicht: (2025)
von: Liang, Zida, et al.
Veröffentlicht: (2025)
Transformation of audio embeddings into interpretable, concept-based representations
von: Zhang, Alice, et al.
Veröffentlicht: (2025)
von: Zhang, Alice, et al.
Veröffentlicht: (2025)
Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs
von: Ye, Charles, et al.
Veröffentlicht: (2026)
von: Ye, Charles, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Weight-based Decomposition: A Case for Bilinear MLPs
von: Pearce, Michael T., et al.
Veröffentlicht: (2024) -
Bilinear autoencoders find interpretable manifolds
von: Dooms, Thomas, et al.
Veröffentlicht: (2026) -
Converting MLPs into Polynomials in Closed Form
von: Belrose, Nora, et al.
Veröffentlicht: (2025) -
Finding Manifolds With Bilinear Autoencoders
von: Dooms, Thomas, et al.
Veröffentlicht: (2025) -
Bilinear Convolution Decomposition for Causal RL Interpretability
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2024)