Divine Benevolence is an $x^2$: GLUs scale asymptotically faster than MLPs
Fuente:
arXiv
Guardado en:
| Autor principal: | Queiruga, Alejandro Francisco |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Interpretability and Generalization Bounds for Learning Spatial Physics
por: Queiruga, Alejandro Francisco, et al.
Publicado: (2025)
por: Queiruga, Alejandro Francisco, et al.
Publicado: (2025)
Estimating the expected output of wide random MLPs more efficiently than sampling
por: Wu, Wilson, et al.
Publicado: (2026)
por: Wu, Wilson, et al.
Publicado: (2026)
Interpolated-MLPs: Controllable Inductive Bias
por: Wu, Sean, et al.
Publicado: (2024)
por: Wu, Sean, et al.
Publicado: (2024)
Converting MLPs into Polynomials in Closed Form
por: Belrose, Nora, et al.
Publicado: (2025)
por: Belrose, Nora, et al.
Publicado: (2025)
MLPs at the EOC: Spectrum of the NTK
por: Terjék, Dávid, et al.
Publicado: (2025)
por: Terjék, Dávid, et al.
Publicado: (2025)
MLPs at the EOC: Concentration of the NTK
por: Terjék, Dávid, et al.
Publicado: (2025)
por: Terjék, Dávid, et al.
Publicado: (2025)
Benchmarking Optimizers for MLPs in Tabular Deep Learning
por: Gorishniy, Yury, et al.
Publicado: (2026)
por: Gorishniy, Yury, et al.
Publicado: (2026)
Benevolent Dictators? On LLM Agent Behavior in Dictator Games
por: Einwiller, Andreas, et al.
Publicado: (2025)
por: Einwiller, Andreas, et al.
Publicado: (2025)
Bilinear MLPs enable weight-based mechanistic interpretability
por: Pearce, Michael T., et al.
Publicado: (2024)
por: Pearce, Michael T., et al.
Publicado: (2024)
Hyperparameter Tuning MLPs for Probabilistic Time Series Forecasting
por: Madhusudhanan, Kiran, et al.
Publicado: (2024)
por: Madhusudhanan, Kiran, et al.
Publicado: (2024)
Mimetic Initialization of MLPs
por: Trockman, Asher, et al.
Publicado: (2026)
por: Trockman, Asher, et al.
Publicado: (2026)
Constructing Efficient Fact-Storing MLPs for Transformers
por: Dugan, Owen, et al.
Publicado: (2025)
por: Dugan, Owen, et al.
Publicado: (2025)
Structural Disentanglement in Bilinear MLPs via Architectural Inductive Bias
por: Nema, Ojasva, et al.
Publicado: (2026)
por: Nema, Ojasva, et al.
Publicado: (2026)
MLPs Learn In-Context on Regression and Classification Tasks
por: Tong, William L., et al.
Publicado: (2024)
por: Tong, William L., et al.
Publicado: (2024)
Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPs
por: Chowdhury, Tanya, et al.
Publicado: (2025)
por: Chowdhury, Tanya, et al.
Publicado: (2025)
Why Attention Fails: The Degeneration of Transformers into MLPs in Time Series Forecasting
por: Liang, Zida, et al.
Publicado: (2025)
por: Liang, Zida, et al.
Publicado: (2025)
Weight-based Decomposition: A Case for Bilinear MLPs
por: Pearce, Michael T., et al.
Publicado: (2024)
por: Pearce, Michael T., et al.
Publicado: (2024)
Training MLPs on Graphs without Supervision
por: Wang, Zehong, et al.
Publicado: (2024)
por: Wang, Zehong, et al.
Publicado: (2024)
How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
por: Demir, Samet, et al.
Publicado: (2025)
por: Demir, Samet, et al.
Publicado: (2025)
Boosting MLPs with a Coarsening Strategy for Long-Term Time Series Forecasting
por: Bian, Nannan, et al.
Publicado: (2024)
por: Bian, Nannan, et al.
Publicado: (2024)
Better by Default: Strong Pre-Tuned MLPs and Boosted Trees on Tabular Data
por: Holzmüller, David, et al.
Publicado: (2024)
por: Holzmüller, David, et al.
Publicado: (2024)
Teaching MLPs to Master Heterogeneous Graph-Structured Knowledge for Efficient and Accurate Inference
por: Liu, Yunhui, et al.
Publicado: (2024)
por: Liu, Yunhui, et al.
Publicado: (2024)
Bayesian Inference with Shaped Deep Non-linear MLPs
por: Hanin, Boris, et al.
Publicado: (2026)
por: Hanin, Boris, et al.
Publicado: (2026)
Emergent Low-Rank Training Dynamics in MLPs with Smooth Activations
por: Xu, Alec S., et al.
Publicado: (2026)
por: Xu, Alec S., et al.
Publicado: (2026)
Channel-Wise MLPs Improve the Generalization of Recurrent Convolutional Networks
por: Breslow, Nathan
Publicado: (2025)
por: Breslow, Nathan
Publicado: (2025)
LightHGNN: Distilling Hypergraph Neural Networks into MLPs for $100\times$ Faster Inference
por: Feng, Yifan, et al.
Publicado: (2024)
por: Feng, Yifan, et al.
Publicado: (2024)
TabICLv2: A better, faster, scalable, and open tabular foundation model
por: Qu, Jingang, et al.
Publicado: (2026)
por: Qu, Jingang, et al.
Publicado: (2026)
From Latent Space to Training Data: Explainable Specialization in Minimal MLPs
por: Alba, Enrique, et al.
Publicado: (2026)
por: Alba, Enrique, et al.
Publicado: (2026)
Provable FDR Control for Deep Feature Selection: Deep MLPs and Beyond
por: Sawaya, Kazuma
Publicado: (2025)
por: Sawaya, Kazuma
Publicado: (2025)
Heuristic Methods are Good Teachers to Distill MLPs for Graph Link Prediction
por: Qin, Zongyue, et al.
Publicado: (2025)
por: Qin, Zongyue, et al.
Publicado: (2025)
Diffusion-Assisted Distillation for Self-Supervised Graph Representation Learning with MLPs
por: Ahn, Seong Jin, et al.
Publicado: (2025)
por: Ahn, Seong Jin, et al.
Publicado: (2025)
TINED: GNNs-to-MLPs by Teacher Injection and Dirichlet Energy Distillation
por: Zhou, Ziang, et al.
Publicado: (2024)
por: Zhou, Ziang, et al.
Publicado: (2024)
Joint control variate for faster black-box variational inference
por: Wang, Xi, et al.
Publicado: (2022)
por: Wang, Xi, et al.
Publicado: (2022)
Team up GBDTs and DNNs: Advancing Efficient and Effective Tabular Prediction with Tree-hybrid MLPs
por: Yan, Jiahuan, et al.
Publicado: (2024)
por: Yan, Jiahuan, et al.
Publicado: (2024)
SimMLP: Training MLPs on Graphs without Supervision
por: Wang, Zehong, et al.
Publicado: (2024)
por: Wang, Zehong, et al.
Publicado: (2024)
Feasibility Study of CNNs and MLPs for Radiation Heat Transfer in 2-D Furnaces with Spectrally Participative Gases
por: TahmasebiMoradi, Axel, et al.
Publicado: (2025)
por: TahmasebiMoradi, Axel, et al.
Publicado: (2025)
Forward-Learned Discrete Diffusion: Learning how to noise to denoise faster
por: Bartosh, Grigory, et al.
Publicado: (2026)
por: Bartosh, Grigory, et al.
Publicado: (2026)
Towards smaller, faster decoder-only transformers: Architectural variants and their implications
por: Suresh, Sathya Krishnan, et al.
Publicado: (2024)
por: Suresh, Sathya Krishnan, et al.
Publicado: (2024)
MLPs and KANs for data-driven learning in physical problems: A performance comparison
por: Pant, Raghav, et al.
Publicado: (2025)
por: Pant, Raghav, et al.
Publicado: (2025)
Latent Algorithmic Structure Precedes Grokking: A Mechanistic Study of ReLU MLPs on Modular Arithmetic
por: Swaroop, Anand
Publicado: (2026)
por: Swaroop, Anand
Publicado: (2026)
Ejemplares similares
-
Interpretability and Generalization Bounds for Learning Spatial Physics
por: Queiruga, Alejandro Francisco, et al.
Publicado: (2025) -
Estimating the expected output of wide random MLPs more efficiently than sampling
por: Wu, Wilson, et al.
Publicado: (2026) -
Interpolated-MLPs: Controllable Inductive Bias
por: Wu, Sean, et al.
Publicado: (2024) -
Converting MLPs into Polynomials in Closed Form
por: Belrose, Nora, et al.
Publicado: (2025) -
MLPs at the EOC: Spectrum of the NTK
por: Terjék, Dávid, et al.
Publicado: (2025)