Understanding MLP-Mixer as a Wide and Sparse MLP
Fuente:
arXiv
Guardado en:
| Autores principales: | Hayase, Tomohiro, Karakida, Ryo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Hierarchical Associative Memory, Parallelized MLP-Mixer, and Symmetry Breaking
por: Karakida, Ryo, et al.
Publicado: (2024)
por: Karakida, Ryo, et al.
Publicado: (2024)
A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention
por: Hayase, Tomohiro, et al.
Publicado: (2026)
por: Hayase, Tomohiro, et al.
Publicado: (2026)
Gaussian Equivalence for Self-Attention: Asymptotic Spectral Analysis of Attention Matrix
por: Hayase, Tomohiro, et al.
Publicado: (2025)
por: Hayase, Tomohiro, et al.
Publicado: (2025)
PreMixer: MLP-Based Pre-training Enhanced MLP-Mixers for Large-scale Traffic Forecasting
por: Zhang, Tongtong, et al.
Publicado: (2024)
por: Zhang, Tongtong, et al.
Publicado: (2024)
MixerFlow: MLP-Mixer meets Normalising Flows
por: English, Eshant, et al.
Publicado: (2023)
por: English, Eshant, et al.
Publicado: (2023)
Fast Jet Tagging with MLP-Mixers on FPGAs
por: Sun, Chang, et al.
Publicado: (2025)
por: Sun, Chang, et al.
Publicado: (2025)
Temporal Graph MLP Mixer for Spatio-Temporal Forecasting
por: Bilal, Muhammad, et al.
Publicado: (2025)
por: Bilal, Muhammad, et al.
Publicado: (2025)
Contextualizing MLP-Mixers Spatiotemporally for Urban Data Forecast at Scale
por: Nie, Tong, et al.
Publicado: (2023)
por: Nie, Tong, et al.
Publicado: (2023)
MLP-SRGAN: A Single-Dimension Super Resolution GAN using MLP-Mixer
por: Mitha, Samir, et al.
Publicado: (2023)
por: Mitha, Samir, et al.
Publicado: (2023)
HSTMixer: A Hierarchical MLP-Mixer for Large-Scale Traffic Forecasting
por: Wang, Yongyao, et al.
Publicado: (2025)
por: Wang, Yongyao, et al.
Publicado: (2025)
A Multi-Scale Decomposition MLP-Mixer for Time Series Analysis
por: Zhong, Shuhan, et al.
Publicado: (2023)
por: Zhong, Shuhan, et al.
Publicado: (2023)
Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models
por: Badger, Benjamin L., et al.
Publicado: (2026)
por: Badger, Benjamin L., et al.
Publicado: (2026)
TSKANMixer: Kolmogorov-Arnold Networks with MLP-Mixer Model for Time Series Forecasting
por: Hong, Young-Chae, et al.
Publicado: (2025)
por: Hong, Young-Chae, et al.
Publicado: (2025)
iMixer: hierarchical Hopfield network implies an invertible, implicit and iterative MLP-Mixer
por: Ota, Toshihiro, et al.
Publicado: (2023)
por: Ota, Toshihiro, et al.
Publicado: (2023)
SEMixer: Semantics Enhanced MLP-Mixer for Multiscale Mixing and Long-term Time Series Forecasting
por: Zhang, Xu, et al.
Publicado: (2026)
por: Zhang, Xu, et al.
Publicado: (2026)
PatchAD: A Lightweight Patch-based MLP-Mixer for Time Series Anomaly Detection
por: Zhong, Zhijie, et al.
Publicado: (2024)
por: Zhong, Zhijie, et al.
Publicado: (2024)
From MLP to NeoMLP: Leveraging Self-Attention for Neural Fields
por: Kofinas, Miltiadis, et al.
Publicado: (2024)
por: Kofinas, Miltiadis, et al.
Publicado: (2024)
BumpNet: A Sparse MLP Framework for Learning PDE Solutions
por: Chiu, Shao-Ting, et al.
Publicado: (2025)
por: Chiu, Shao-Ting, et al.
Publicado: (2025)
On the Parameterization of Second-Order Optimization Effective Towards the Infinite Width
por: Ishikawa, Satoki, et al.
Publicado: (2023)
por: Ishikawa, Satoki, et al.
Publicado: (2023)
Rethinking the shape convention of an MLP
por: Chen, Meng-Hsi, et al.
Publicado: (2025)
por: Chen, Meng-Hsi, et al.
Publicado: (2025)
Spatial Transfer Learning with Simple MLP
por: Yang, Hongjian
Publicado: (2024)
por: Yang, Hongjian
Publicado: (2024)
PowerMLP: An Efficient Version of KAN
por: Qiu, Ruichen, et al.
Publicado: (2024)
por: Qiu, Ruichen, et al.
Publicado: (2024)
KAN or MLP: A Fairer Comparison
por: Yu, Runpeng, et al.
Publicado: (2024)
por: Yu, Runpeng, et al.
Publicado: (2024)
Optimal Layer Selection for Latent Data Augmentation
por: Takase, Tomoumi, et al.
Publicado: (2024)
por: Takase, Tomoumi, et al.
Publicado: (2024)
GraphMLP: A Graph MLP-Like Architecture for 3D Human Pose Estimation
por: Li, Wenhao, et al.
Publicado: (2022)
por: Li, Wenhao, et al.
Publicado: (2022)
Infinite-Width Limit of a Single Attention Layer: Analysis via Tensor Programs
por: Sakai, Mana, et al.
Publicado: (2025)
por: Sakai, Mana, et al.
Publicado: (2025)
Maintaining and Managing Road Quality:Using MLP and DNN
por: Maotwana, Makgotso Jacqueline
Publicado: (2024)
por: Maotwana, Makgotso Jacqueline
Publicado: (2024)
An All-MLP Sequence Modeling Architecture That Excels at Copying
por: Cui, Chenwei, et al.
Publicado: (2024)
por: Cui, Chenwei, et al.
Publicado: (2024)
Can an MLP Absorb Its Own Skip Connection?
por: Mijoski, Antonij, et al.
Publicado: (2026)
por: Mijoski, Antonij, et al.
Publicado: (2026)
Efficient LLMs with AMP: Attention Heads and MLP Pruning
por: Mugnaini, Leandro Giusti, et al.
Publicado: (2025)
por: Mugnaini, Leandro Giusti, et al.
Publicado: (2025)
End to End Autoencoder MLP Framework for Sepsis Prediction
por: Cai, Hejiang, et al.
Publicado: (2025)
por: Cai, Hejiang, et al.
Publicado: (2025)
KAN v.s. MLP for Offline Reinforcement Learning
por: Guo, Haihong, et al.
Publicado: (2024)
por: Guo, Haihong, et al.
Publicado: (2024)
Self-attention Networks Localize When QK-eigenspectrum Concentrates
por: Bao, Han, et al.
Publicado: (2024)
por: Bao, Han, et al.
Publicado: (2024)
Continual Learning in Modern Hopfield Networks with an Application to Diffusion Models
por: Takeda, Ken, et al.
Publicado: (2026)
por: Takeda, Ken, et al.
Publicado: (2026)
D2-MLP: Dynamic Decomposed MLP Mixer for Medical Image Segmentation
por: Yang, Jin, et al.
Publicado: (2024)
por: Yang, Jin, et al.
Publicado: (2024)
A Multimodal Fusion Model Leveraging MLP Mixer and Handcrafted Features-based Deep Learning Networks for Facial Palsy Detection
por: Oo, Heng Yim Nicole, et al.
Publicado: (2025)
por: Oo, Heng Yim Nicole, et al.
Publicado: (2025)
Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians
por: Tomihari, Akiyoshi, et al.
Publicado: (2025)
por: Tomihari, Akiyoshi, et al.
Publicado: (2025)
HyperMLP: An Integrated Perspective for Sequence Modeling
por: Lu, Jiecheng, et al.
Publicado: (2026)
por: Lu, Jiecheng, et al.
Publicado: (2026)
Hypergraph-MLP: Learning on Hypergraphs without Message Passing
por: Tang, Bohan, et al.
Publicado: (2023)
por: Tang, Bohan, et al.
Publicado: (2023)
MLP-KAN: Unifying Deep Representation and Function Learning
por: He, Yunhong, et al.
Publicado: (2024)
por: He, Yunhong, et al.
Publicado: (2024)
Ejemplares similares
-
Hierarchical Associative Memory, Parallelized MLP-Mixer, and Symmetry Breaking
por: Karakida, Ryo, et al.
Publicado: (2024) -
A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention
por: Hayase, Tomohiro, et al.
Publicado: (2026) -
Gaussian Equivalence for Self-Attention: Asymptotic Spectral Analysis of Attention Matrix
por: Hayase, Tomohiro, et al.
Publicado: (2025) -
PreMixer: MLP-Based Pre-training Enhanced MLP-Mixers for Large-scale Traffic Forecasting
por: Zhang, Tongtong, et al.
Publicado: (2024) -
MixerFlow: MLP-Mixer meets Normalising Flows
por: English, Eshant, et al.
Publicado: (2023)