Rethinking the shape convention of an MLP
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Meng-Hsi, Lee, Yu-Ang, Liao, Feng-Ting, Shiu, Da-shan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revisiting the Shape Convention of Transformer Language Models
by: Liao, Feng-Ting, et al.
Published: (2026)
by: Liao, Feng-Ting, et al.
Published: (2026)
Latent Flow Transformer
by: Wu, Yen-Chen, et al.
Published: (2025)
by: Wu, Yen-Chen, et al.
Published: (2025)
Exact, Tractable Gauss-Newton Optimization in Deep Reversible Architectures Reveal Poor Generalization
by: Buffelli, Davide, et al.
Published: (2024)
by: Buffelli, Davide, et al.
Published: (2024)
KAN or MLP: A Fairer Comparison
by: Yu, Runpeng, et al.
Published: (2024)
by: Yu, Runpeng, et al.
Published: (2024)
Teach Harder, Learn Poorer: Rethinking Hard Sample Distillation for GNN-to-MLP Knowledge Distillation
by: Wu, Lirong, et al.
Published: (2024)
by: Wu, Lirong, et al.
Published: (2024)
From MLP to NeoMLP: Leveraging Self-Attention for Neural Fields
by: Kofinas, Miltiadis, et al.
Published: (2024)
by: Kofinas, Miltiadis, et al.
Published: (2024)
Towards a Foundation Model for Communication Systems
by: Buffelli, Davide, et al.
Published: (2025)
by: Buffelli, Davide, et al.
Published: (2025)
Rethinking Refinement: Correcting Generative Bias without Noise Injection
by: Peng, Xin, et al.
Published: (2026)
by: Peng, Xin, et al.
Published: (2026)
FedMLP: Federated Multi-Label Medical Image Classification under Task Heterogeneity
by: Sun, Zhaobin, et al.
Published: (2024)
by: Sun, Zhaobin, et al.
Published: (2024)
KAN v.s. MLP for Offline Reinforcement Learning
by: Guo, Haihong, et al.
Published: (2024)
by: Guo, Haihong, et al.
Published: (2024)
TSKANMixer: Kolmogorov-Arnold Networks with MLP-Mixer Model for Time Series Forecasting
by: Hong, Young-Chae, et al.
Published: (2025)
by: Hong, Young-Chae, et al.
Published: (2025)
(GG) MoE vs. MLP on Tabular Data
by: Chernov, Andrei
Published: (2025)
by: Chernov, Andrei
Published: (2025)
Aggregation-aware MLP: An Unsupervised Approach for Graph Message-passing
by: Xie, Xuanting, et al.
Published: (2025)
by: Xie, Xuanting, et al.
Published: (2025)
AdaGMLP: AdaBoosting GNN-to-MLP Knowledge Distillation
by: Lu, Weigang, et al.
Published: (2024)
by: Lu, Weigang, et al.
Published: (2024)
A Multi-Scale Decomposition MLP-Mixer for Time Series Analysis
by: Zhong, Shuhan, et al.
Published: (2023)
by: Zhong, Shuhan, et al.
Published: (2023)
Rethinking the Representation in Federated Unsupervised Learning with Non-IID Data
by: Liao, Xinting, et al.
Published: (2024)
by: Liao, Xinting, et al.
Published: (2024)
XLinear: Frequency-Enhanced MLP with CrossFilter for Robust Long-Range Forecasting
by: Ao, Xiang
Published: (2026)
by: Ao, Xiang
Published: (2026)
HyperMLP: An Integrated Perspective for Sequence Modeling
by: Lu, Jiecheng, et al.
Published: (2026)
by: Lu, Jiecheng, et al.
Published: (2026)
Learning to Select MCP Algorithms: From Traditional ML to Dual-Channel GAT-MLP
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Unlocking the Power of Patch: Patch-Based MLP for Long-Term Time Series Forecasting
by: Tang, Peiwang, et al.
Published: (2024)
by: Tang, Peiwang, et al.
Published: (2024)
Causal-Driven Feature Evaluation for Cross-Domain Image Classification
by: Cheng, Chen, et al.
Published: (2026)
by: Cheng, Chen, et al.
Published: (2026)
Teaching MLP More Graph Information: A Three-stage Multitask Knowledge Distillation Framework
by: Li, Junxian, et al.
Published: (2024)
by: Li, Junxian, et al.
Published: (2024)
Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models
by: Huang, Wei-Ping, et al.
Published: (2026)
by: Huang, Wei-Ping, et al.
Published: (2026)
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
by: Tian, Yuandong, et al.
Published: (2023)
by: Tian, Yuandong, et al.
Published: (2023)
Rethinking Constraint Awareness for Efficient State Embedding of Neural Routing Solver
by: Yu, Canhong, et al.
Published: (2026)
by: Yu, Canhong, et al.
Published: (2026)
SwinGNN: Rethinking Permutation Invariance in Diffusion Models for Graph Generation
by: Yan, Qi, et al.
Published: (2023)
by: Yan, Qi, et al.
Published: (2023)
GraphMLP: A Graph MLP-Like Architecture for 3D Human Pose Estimation
by: Li, Wenhao, et al.
Published: (2022)
by: Li, Wenhao, et al.
Published: (2022)
Rethinking Efficient Graph Coarsening via a Non-Selfishness Principle
by: Bai, Xu, et al.
Published: (2026)
by: Bai, Xu, et al.
Published: (2026)
M3-Net: A Cost-Effective Graph-Free MLP-Based Model for Traffic Prediction
by: Jin, Guangyin, et al.
Published: (2025)
by: Jin, Guangyin, et al.
Published: (2025)
Rethinking industrial artificial intelligence: a unified foundation framework
by: Lee, Jay, et al.
Published: (2025)
by: Lee, Jay, et al.
Published: (2025)
DynamicGate MLP Conditional Computation via Learned Structural Dropout and Input Dependent Gating for Functional Plasticity
by: Choi, Yong Il
Published: (2026)
by: Choi, Yong Il
Published: (2026)
Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level Granularity
by: Hsu, Chan-Jan, et al.
Published: (2025)
by: Hsu, Chan-Jan, et al.
Published: (2025)
KAN versus MLP on Irregular or Noisy Functions
by: Zeng, Chen, et al.
Published: (2024)
by: Zeng, Chen, et al.
Published: (2024)
Bridging KAN and MLP: MJKAN, a Hybrid Architecture with Both Efficiency and Expressiveness
by: Joo, Hanseon, et al.
Published: (2025)
by: Joo, Hanseon, et al.
Published: (2025)
Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction
by: Ding, Fei, et al.
Published: (2026)
by: Ding, Fei, et al.
Published: (2026)
SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models
by: Zhu, Hourun, et al.
Published: (2025)
by: Zhu, Hourun, et al.
Published: (2025)
Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical Reasoning
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models
by: Badger, Benjamin L., et al.
Published: (2026)
by: Badger, Benjamin L., et al.
Published: (2026)
Incorporating Exponential Smoothing into MLP: A Simple but Effective Sequence Model
by: Chu, Jiqun, et al.
Published: (2024)
by: Chu, Jiqun, et al.
Published: (2024)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
Similar Items
-
Revisiting the Shape Convention of Transformer Language Models
by: Liao, Feng-Ting, et al.
Published: (2026) -
Latent Flow Transformer
by: Wu, Yen-Chen, et al.
Published: (2025) -
Exact, Tractable Gauss-Newton Optimization in Deep Reversible Architectures Reveal Poor Generalization
by: Buffelli, Davide, et al.
Published: (2024) -
KAN or MLP: A Fairer Comparison
by: Yu, Runpeng, et al.
Published: (2024) -
Teach Harder, Learn Poorer: Rethinking Hard Sample Distillation for GNN-to-MLP Knowledge Distillation
by: Wu, Lirong, et al.
Published: (2024)