Converting MLPs into Polynomials in Closed Form
Fuente:
arXiv
Saved in:
| Main Authors: | Belrose, Nora, Rigg, Alice |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Weight-based Decomposition: A Case for Bilinear MLPs
by: Pearce, Michael T., et al.
Published: (2024)
by: Pearce, Michael T., et al.
Published: (2024)
Bilinear MLPs enable weight-based mechanistic interpretability
by: Pearce, Michael T., et al.
Published: (2024)
by: Pearce, Michael T., et al.
Published: (2024)
Estimating the Probability of Sampling a Trained Neural Network at Random
by: Scherlis, Adam, et al.
Published: (2025)
by: Scherlis, Adam, et al.
Published: (2025)
Evaluating SAE interpretability without explanations
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Slowing Learning by Erasing Simple Features
by: Quirke, Lucia, et al.
Published: (2025)
by: Quirke, Lucia, et al.
Published: (2025)
Sparse Autoencoders Trained on the Same Data Learn Different Features
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Balancing Label Quantity and Quality for Scalable Elicitation
by: Mallen, Alex, et al.
Published: (2024)
by: Mallen, Alex, et al.
Published: (2024)
Understanding Gradient Descent through the Training Jacobian
by: Belrose, Nora, et al.
Published: (2024)
by: Belrose, Nora, et al.
Published: (2024)
Partially Rewriting a Transformer in Natural Language
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Examining Two Hop Reasoning Through Information Content Scaling
by: Johnston, David, et al.
Published: (2025)
by: Johnston, David, et al.
Published: (2025)
Binary Sparse Coding for Interpretability
by: Quirke, Lucia, et al.
Published: (2025)
by: Quirke, Lucia, et al.
Published: (2025)
Transcoders Beat Sparse Autoencoders for Interpretability
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Refusal in LLMs is an Affine Function
by: Marshall, Thomas, et al.
Published: (2024)
by: Marshall, Thomas, et al.
Published: (2024)
Does Transformer Interpretability Transfer to RNNs?
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
Mechanistic Anomaly Detection for "Quirky" Language Models
by: Johnston, David O., et al.
Published: (2025)
by: Johnston, David O., et al.
Published: (2025)
Bilinear Convolution Decomposition for Causal RL Interpretability
by: Oozeer, Narmeen, et al.
Published: (2024)
by: Oozeer, Narmeen, et al.
Published: (2024)
Automatically Interpreting Millions of Features in Large Language Models
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
Neural Networks Learn Statistics of Increasing Complexity
by: Belrose, Nora, et al.
Published: (2024)
by: Belrose, Nora, et al.
Published: (2024)
Eliciting Latent Knowledge from Quirky Language Models
by: Mallen, Alex, et al.
Published: (2023)
by: Mallen, Alex, et al.
Published: (2023)
Distribution-Aware Feature Selection for SAEs
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Converting Transformers into DGNNs Form
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
Interpolated-MLPs: Controllable Inductive Bias
by: Wu, Sean, et al.
Published: (2024)
by: Wu, Sean, et al.
Published: (2024)
MLPs at the EOC: Spectrum of the NTK
by: Terjék, Dávid, et al.
Published: (2025)
by: Terjék, Dávid, et al.
Published: (2025)
MLPs at the EOC: Concentration of the NTK
by: Terjék, Dávid, et al.
Published: (2025)
by: Terjék, Dávid, et al.
Published: (2025)
Benchmarking Optimizers for MLPs in Tabular Deep Learning
by: Gorishniy, Yury, et al.
Published: (2026)
by: Gorishniy, Yury, et al.
Published: (2026)
Detecting and Characterizing Planning in Language Models
by: Nainani, Jatin, et al.
Published: (2025)
by: Nainani, Jatin, et al.
Published: (2025)
Hyperparameter Tuning MLPs for Probabilistic Time Series Forecasting
by: Madhusudhanan, Kiran, et al.
Published: (2024)
by: Madhusudhanan, Kiran, et al.
Published: (2024)
Closed-Form Diffusion Models
by: Scarvelis, Christopher, et al.
Published: (2023)
by: Scarvelis, Christopher, et al.
Published: (2023)
Mimetic Initialization of MLPs
by: Trockman, Asher, et al.
Published: (2026)
by: Trockman, Asher, et al.
Published: (2026)
Constructing Efficient Fact-Storing MLPs for Transformers
by: Dugan, Owen, et al.
Published: (2025)
by: Dugan, Owen, et al.
Published: (2025)
Structural Disentanglement in Bilinear MLPs via Architectural Inductive Bias
by: Nema, Ojasva, et al.
Published: (2026)
by: Nema, Ojasva, et al.
Published: (2026)
Closed-Form Last Layer Optimization
by: Galashov, Alexandre, et al.
Published: (2025)
by: Galashov, Alexandre, et al.
Published: (2025)
LEACE: Perfect linear concept erasure in closed form
by: Belrose, Nora, et al.
Published: (2023)
by: Belrose, Nora, et al.
Published: (2023)
Eliciting Latent Predictions from Transformers with the Tuned Lens
by: Belrose, Nora, et al.
Published: (2023)
by: Belrose, Nora, et al.
Published: (2023)
MLPs Learn In-Context on Regression and Classification Tasks
by: Tong, William L., et al.
Published: (2024)
by: Tong, William L., et al.
Published: (2024)
Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPs
by: Chowdhury, Tanya, et al.
Published: (2025)
by: Chowdhury, Tanya, et al.
Published: (2025)
Why Attention Fails: The Degeneration of Transformers into MLPs in Time Series Forecasting
by: Liang, Zida, et al.
Published: (2025)
by: Liang, Zida, et al.
Published: (2025)
nn2poly: An R Package for Converting Neural Networks into Interpretable Polynomials
by: Morala, Pablo, et al.
Published: (2024)
by: Morala, Pablo, et al.
Published: (2024)
Training MLPs on Graphs without Supervision
by: Wang, Zehong, et al.
Published: (2024)
by: Wang, Zehong, et al.
Published: (2024)
How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
by: Demir, Samet, et al.
Published: (2025)
by: Demir, Samet, et al.
Published: (2025)
Similar Items
-
Weight-based Decomposition: A Case for Bilinear MLPs
by: Pearce, Michael T., et al.
Published: (2024) -
Bilinear MLPs enable weight-based mechanistic interpretability
by: Pearce, Michael T., et al.
Published: (2024) -
Estimating the Probability of Sampling a Trained Neural Network at Random
by: Scherlis, Adam, et al.
Published: (2025) -
Evaluating SAE interpretability without explanations
by: Paulo, Gonçalo, et al.
Published: (2025) -
Slowing Learning by Erasing Simple Features
by: Quirke, Lucia, et al.
Published: (2025)