Weight-based Decomposition: A Case for Bilinear MLPs
Fuente:
arXiv
Saved in:
| Main Authors: | Pearce, Michael T., Dooms, Thomas, Rigg, Alice |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bilinear MLPs enable weight-based mechanistic interpretability
by: Pearce, Michael T., et al.
Published: (2024)
by: Pearce, Michael T., et al.
Published: (2024)
Bilinear Convolution Decomposition for Causal RL Interpretability
by: Oozeer, Narmeen, et al.
Published: (2024)
by: Oozeer, Narmeen, et al.
Published: (2024)
Converting MLPs into Polynomials in Closed Form
by: Belrose, Nora, et al.
Published: (2025)
by: Belrose, Nora, et al.
Published: (2025)
Finding Manifolds With Bilinear Autoencoders
by: Dooms, Thomas, et al.
Published: (2025)
by: Dooms, Thomas, et al.
Published: (2025)
Preserving Bilinear Weight Spectra with a Signed and Shrunk Quadratic Activation Function
by: Abohwo, Jason, et al.
Published: (2025)
by: Abohwo, Jason, et al.
Published: (2025)
Constructing Efficient Fact-Storing MLPs for Transformers
by: Dugan, Owen, et al.
Published: (2025)
by: Dugan, Owen, et al.
Published: (2025)
Bilinear autoencoders find interpretable manifolds
by: Dooms, Thomas, et al.
Published: (2026)
by: Dooms, Thomas, et al.
Published: (2026)
Emergent Low-Rank Training Dynamics in MLPs with Smooth Activations
by: Xu, Alec S., et al.
Published: (2026)
by: Xu, Alec S., et al.
Published: (2026)
Channel-Wise MLPs Improve the Generalization of Recurrent Convolutional Networks
by: Breslow, Nathan
Published: (2025)
by: Breslow, Nathan
Published: (2025)
Mimetic Initialization of MLPs
by: Trockman, Asher, et al.
Published: (2026)
by: Trockman, Asher, et al.
Published: (2026)
From Latent Space to Training Data: Explainable Specialization in Minimal MLPs
by: Alba, Enrique, et al.
Published: (2026)
by: Alba, Enrique, et al.
Published: (2026)
Heuristic Methods are Good Teachers to Distill MLPs for Graph Link Prediction
by: Qin, Zongyue, et al.
Published: (2025)
by: Qin, Zongyue, et al.
Published: (2025)
Diffusion-Assisted Distillation for Self-Supervised Graph Representation Learning with MLPs
by: Ahn, Seong Jin, et al.
Published: (2025)
by: Ahn, Seong Jin, et al.
Published: (2025)
BECAUSE: Bilinear Causal Representation for Generalizable Offline Model-based Reinforcement Learning
by: Lin, Haohong, et al.
Published: (2024)
by: Lin, Haohong, et al.
Published: (2024)
Training MLPs on Graphs without Supervision
by: Wang, Zehong, et al.
Published: (2024)
by: Wang, Zehong, et al.
Published: (2024)
Modular addition without black-boxes: Compressing explanations of MLPs that compute numerical integration
by: Yip, Chun Hei, et al.
Published: (2024)
by: Yip, Chun Hei, et al.
Published: (2024)
Learning to Model Graph Structural Information on MLPs via Graph Structure Self-Contrasting
by: Wu, Lirong, et al.
Published: (2024)
by: Wu, Lirong, et al.
Published: (2024)
Edge-free but Structure-aware: Prototype-Guided Knowledge Distillation from GNNs to MLPs
by: Wu, Taiqiang, et al.
Published: (2023)
by: Wu, Taiqiang, et al.
Published: (2023)
MDMLP-EIA: Multi-domain Dynamic MLPs with Energy Invariant Attention for Time Series Forecasting
by: Zhang, Hu, et al.
Published: (2025)
by: Zhang, Hu, et al.
Published: (2025)
MLPMoE: Zero-Shot Architectural Metamorphosis of Dense LLM MLPs into Static Mixture-of-Experts
by: Novikov, Ivan
Published: (2025)
by: Novikov, Ivan
Published: (2025)
Beyond Johnson-Lindenstrauss: Uniform Bounds for Sketched Bilinear Forms
by: Deb, Rohan, et al.
Published: (2025)
by: Deb, Rohan, et al.
Published: (2025)
SimMLP: Training MLPs on Graphs without Supervision
by: Wang, Zehong, et al.
Published: (2024)
by: Wang, Zehong, et al.
Published: (2024)
Network-Aware Bilinear Tokenization for Brain Functional Connectivity Representation Learning
by: Milecki, Leo, et al.
Published: (2026)
by: Milecki, Leo, et al.
Published: (2026)
Bilinear representation mitigates reversal curse and enables consistent model editing
by: Kim, Dong-Kyum, et al.
Published: (2025)
by: Kim, Dong-Kyum, et al.
Published: (2025)
MIBP-Cert: Certified Training against Data Perturbations with Mixed-Integer Bilinear Programs
by: Lorenz, Tobias, et al.
Published: (2024)
by: Lorenz, Tobias, et al.
Published: (2024)
Looped ReLU MLPs May Be All You Need as Practical Programmable Computers
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
Structural Disentanglement in Bilinear MLPs via Architectural Inductive Bias
by: Nema, Ojasva, et al.
Published: (2026)
by: Nema, Ojasva, et al.
Published: (2026)
VQGraph: Rethinking Graph Representation Space for Bridging GNNs and MLPs
by: Yang, Ling, et al.
Published: (2023)
by: Yang, Ling, et al.
Published: (2023)
LORA-CRAFT: Cross-layer Rank Adaptation via Frozen Tucker Decomposition of Pre-trained Attention Weights
by: Dewage, Kasun, et al.
Published: (2026)
by: Dewage, Kasun, et al.
Published: (2026)
Estimating Worst-Case Frontier Risks of Open-Weight LLMs
by: Wallace, Eric, et al.
Published: (2025)
by: Wallace, Eric, et al.
Published: (2025)
Attention-based Iterative Decomposition for Tensor Product Representation
by: Park, Taewon, et al.
Published: (2024)
by: Park, Taewon, et al.
Published: (2024)
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
by: Lau, Tim Tsz-Kit, et al.
Published: (2026)
by: Lau, Tim Tsz-Kit, et al.
Published: (2026)
ADMIRE-BayesOpt: Accelerated Data MIxture RE-weighting for Language Models with Bayesian Optimization
by: Chen, Shengzhuang, et al.
Published: (2025)
by: Chen, Shengzhuang, et al.
Published: (2025)
Discrete Dictionary-based Decomposition Layer for Structured Representation Learning
by: Park, Taewon, et al.
Published: (2024)
by: Park, Taewon, et al.
Published: (2024)
SamBaTen: Sampling-based Batch Incremental Tensor Decomposition
by: Gujral, Ekta, et al.
Published: (2017)
by: Gujral, Ekta, et al.
Published: (2017)
Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
by: Cho, Yoonjun, et al.
Published: (2025)
by: Cho, Yoonjun, et al.
Published: (2025)
On the Weight Dynamics of Deep Normalized Networks
by: Mehmeti-Göpel, Christian H. X. Ali, et al.
Published: (2023)
by: Mehmeti-Göpel, Christian H. X. Ali, et al.
Published: (2023)
HDNet: Physics-Inspired Neural Network for Flow Estimation based on Helmholtz Decomposition
by: Qi, Miao, et al.
Published: (2024)
by: Qi, Miao, et al.
Published: (2024)
Decomposition-based multi-scale transformer framework for time series anomaly detection
by: Zhang, Wenxin, et al.
Published: (2025)
by: Zhang, Wenxin, et al.
Published: (2025)
Optimal Signal Decomposition-based Multi-Stage Learning for Battery Health Estimation
by: Pamshetti, Vijay Babu, et al.
Published: (2025)
by: Pamshetti, Vijay Babu, et al.
Published: (2025)
Similar Items
-
Bilinear MLPs enable weight-based mechanistic interpretability
by: Pearce, Michael T., et al.
Published: (2024) -
Bilinear Convolution Decomposition for Causal RL Interpretability
by: Oozeer, Narmeen, et al.
Published: (2024) -
Converting MLPs into Polynomials in Closed Form
by: Belrose, Nora, et al.
Published: (2025) -
Finding Manifolds With Bilinear Autoencoders
by: Dooms, Thomas, et al.
Published: (2025) -
Preserving Bilinear Weight Spectra with a Signed and Shrunk Quadratic Activation Function
by: Abohwo, Jason, et al.
Published: (2025)