Divine Benevolence is an $x^2$: GLUs scale asymptotically faster than MLPs
Fuente:
arXiv
Salvato in:
| Autore principale: | Queiruga, Alejandro Francisco |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Interpretability and Generalization Bounds for Learning Spatial Physics
di: Queiruga, Alejandro Francisco, et al.
Pubblicazione: (2025)
di: Queiruga, Alejandro Francisco, et al.
Pubblicazione: (2025)
Estimating the expected output of wide random MLPs more efficiently than sampling
di: Wu, Wilson, et al.
Pubblicazione: (2026)
di: Wu, Wilson, et al.
Pubblicazione: (2026)
Interpolated-MLPs: Controllable Inductive Bias
di: Wu, Sean, et al.
Pubblicazione: (2024)
di: Wu, Sean, et al.
Pubblicazione: (2024)
Converting MLPs into Polynomials in Closed Form
di: Belrose, Nora, et al.
Pubblicazione: (2025)
di: Belrose, Nora, et al.
Pubblicazione: (2025)
MLPs at the EOC: Spectrum of the NTK
di: Terjék, Dávid, et al.
Pubblicazione: (2025)
di: Terjék, Dávid, et al.
Pubblicazione: (2025)
MLPs at the EOC: Concentration of the NTK
di: Terjék, Dávid, et al.
Pubblicazione: (2025)
di: Terjék, Dávid, et al.
Pubblicazione: (2025)
Benchmarking Optimizers for MLPs in Tabular Deep Learning
di: Gorishniy, Yury, et al.
Pubblicazione: (2026)
di: Gorishniy, Yury, et al.
Pubblicazione: (2026)
Benevolent Dictators? On LLM Agent Behavior in Dictator Games
di: Einwiller, Andreas, et al.
Pubblicazione: (2025)
di: Einwiller, Andreas, et al.
Pubblicazione: (2025)
Bilinear MLPs enable weight-based mechanistic interpretability
di: Pearce, Michael T., et al.
Pubblicazione: (2024)
di: Pearce, Michael T., et al.
Pubblicazione: (2024)
Hyperparameter Tuning MLPs for Probabilistic Time Series Forecasting
di: Madhusudhanan, Kiran, et al.
Pubblicazione: (2024)
di: Madhusudhanan, Kiran, et al.
Pubblicazione: (2024)
Mimetic Initialization of MLPs
di: Trockman, Asher, et al.
Pubblicazione: (2026)
di: Trockman, Asher, et al.
Pubblicazione: (2026)
Constructing Efficient Fact-Storing MLPs for Transformers
di: Dugan, Owen, et al.
Pubblicazione: (2025)
di: Dugan, Owen, et al.
Pubblicazione: (2025)
Structural Disentanglement in Bilinear MLPs via Architectural Inductive Bias
di: Nema, Ojasva, et al.
Pubblicazione: (2026)
di: Nema, Ojasva, et al.
Pubblicazione: (2026)
MLPs Learn In-Context on Regression and Classification Tasks
di: Tong, William L., et al.
Pubblicazione: (2024)
di: Tong, William L., et al.
Pubblicazione: (2024)
Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPs
di: Chowdhury, Tanya, et al.
Pubblicazione: (2025)
di: Chowdhury, Tanya, et al.
Pubblicazione: (2025)
Why Attention Fails: The Degeneration of Transformers into MLPs in Time Series Forecasting
di: Liang, Zida, et al.
Pubblicazione: (2025)
di: Liang, Zida, et al.
Pubblicazione: (2025)
Weight-based Decomposition: A Case for Bilinear MLPs
di: Pearce, Michael T., et al.
Pubblicazione: (2024)
di: Pearce, Michael T., et al.
Pubblicazione: (2024)
Training MLPs on Graphs without Supervision
di: Wang, Zehong, et al.
Pubblicazione: (2024)
di: Wang, Zehong, et al.
Pubblicazione: (2024)
How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
di: Demir, Samet, et al.
Pubblicazione: (2025)
di: Demir, Samet, et al.
Pubblicazione: (2025)
Boosting MLPs with a Coarsening Strategy for Long-Term Time Series Forecasting
di: Bian, Nannan, et al.
Pubblicazione: (2024)
di: Bian, Nannan, et al.
Pubblicazione: (2024)
Better by Default: Strong Pre-Tuned MLPs and Boosted Trees on Tabular Data
di: Holzmüller, David, et al.
Pubblicazione: (2024)
di: Holzmüller, David, et al.
Pubblicazione: (2024)
Teaching MLPs to Master Heterogeneous Graph-Structured Knowledge for Efficient and Accurate Inference
di: Liu, Yunhui, et al.
Pubblicazione: (2024)
di: Liu, Yunhui, et al.
Pubblicazione: (2024)
Bayesian Inference with Shaped Deep Non-linear MLPs
di: Hanin, Boris, et al.
Pubblicazione: (2026)
di: Hanin, Boris, et al.
Pubblicazione: (2026)
Emergent Low-Rank Training Dynamics in MLPs with Smooth Activations
di: Xu, Alec S., et al.
Pubblicazione: (2026)
di: Xu, Alec S., et al.
Pubblicazione: (2026)
Channel-Wise MLPs Improve the Generalization of Recurrent Convolutional Networks
di: Breslow, Nathan
Pubblicazione: (2025)
di: Breslow, Nathan
Pubblicazione: (2025)
LightHGNN: Distilling Hypergraph Neural Networks into MLPs for $100\times$ Faster Inference
di: Feng, Yifan, et al.
Pubblicazione: (2024)
di: Feng, Yifan, et al.
Pubblicazione: (2024)
TabICLv2: A better, faster, scalable, and open tabular foundation model
di: Qu, Jingang, et al.
Pubblicazione: (2026)
di: Qu, Jingang, et al.
Pubblicazione: (2026)
From Latent Space to Training Data: Explainable Specialization in Minimal MLPs
di: Alba, Enrique, et al.
Pubblicazione: (2026)
di: Alba, Enrique, et al.
Pubblicazione: (2026)
Provable FDR Control for Deep Feature Selection: Deep MLPs and Beyond
di: Sawaya, Kazuma
Pubblicazione: (2025)
di: Sawaya, Kazuma
Pubblicazione: (2025)
Heuristic Methods are Good Teachers to Distill MLPs for Graph Link Prediction
di: Qin, Zongyue, et al.
Pubblicazione: (2025)
di: Qin, Zongyue, et al.
Pubblicazione: (2025)
Diffusion-Assisted Distillation for Self-Supervised Graph Representation Learning with MLPs
di: Ahn, Seong Jin, et al.
Pubblicazione: (2025)
di: Ahn, Seong Jin, et al.
Pubblicazione: (2025)
TINED: GNNs-to-MLPs by Teacher Injection and Dirichlet Energy Distillation
di: Zhou, Ziang, et al.
Pubblicazione: (2024)
di: Zhou, Ziang, et al.
Pubblicazione: (2024)
Joint control variate for faster black-box variational inference
di: Wang, Xi, et al.
Pubblicazione: (2022)
di: Wang, Xi, et al.
Pubblicazione: (2022)
Team up GBDTs and DNNs: Advancing Efficient and Effective Tabular Prediction with Tree-hybrid MLPs
di: Yan, Jiahuan, et al.
Pubblicazione: (2024)
di: Yan, Jiahuan, et al.
Pubblicazione: (2024)
SimMLP: Training MLPs on Graphs without Supervision
di: Wang, Zehong, et al.
Pubblicazione: (2024)
di: Wang, Zehong, et al.
Pubblicazione: (2024)
Feasibility Study of CNNs and MLPs for Radiation Heat Transfer in 2-D Furnaces with Spectrally Participative Gases
di: TahmasebiMoradi, Axel, et al.
Pubblicazione: (2025)
di: TahmasebiMoradi, Axel, et al.
Pubblicazione: (2025)
Forward-Learned Discrete Diffusion: Learning how to noise to denoise faster
di: Bartosh, Grigory, et al.
Pubblicazione: (2026)
di: Bartosh, Grigory, et al.
Pubblicazione: (2026)
Towards smaller, faster decoder-only transformers: Architectural variants and their implications
di: Suresh, Sathya Krishnan, et al.
Pubblicazione: (2024)
di: Suresh, Sathya Krishnan, et al.
Pubblicazione: (2024)
MLPs and KANs for data-driven learning in physical problems: A performance comparison
di: Pant, Raghav, et al.
Pubblicazione: (2025)
di: Pant, Raghav, et al.
Pubblicazione: (2025)
Latent Algorithmic Structure Precedes Grokking: A Mechanistic Study of ReLU MLPs on Modular Arithmetic
di: Swaroop, Anand
Pubblicazione: (2026)
di: Swaroop, Anand
Pubblicazione: (2026)
Documenti analoghi
-
Interpretability and Generalization Bounds for Learning Spatial Physics
di: Queiruga, Alejandro Francisco, et al.
Pubblicazione: (2025) -
Estimating the expected output of wide random MLPs more efficiently than sampling
di: Wu, Wilson, et al.
Pubblicazione: (2026) -
Interpolated-MLPs: Controllable Inductive Bias
di: Wu, Sean, et al.
Pubblicazione: (2024) -
Converting MLPs into Polynomials in Closed Form
di: Belrose, Nora, et al.
Pubblicazione: (2025) -
MLPs at the EOC: Spectrum of the NTK
di: Terjék, Dávid, et al.
Pubblicazione: (2025)