An All-MLP Sequence Modeling Architecture That Excels at Copying
Fuente:
arXiv
Saved in:
| Main Authors: | Cui, Chenwei, Yan, Zehao, Muhawenayo, Gedeon, Kerner, Hannah |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Does the Spatial Distribution of Pre-training Data Affect Geospatial Foundation Models?
by: Purohit, Mirali, et al.
Published: (2025)
by: Purohit, Mirali, et al.
Published: (2025)
Pretrain Where? Investigating How Pretraining Data Diversity Impacts Geospatial Foundation Model Performance
by: Kaur, Amandeep, et al.
Published: (2026)
by: Kaur, Amandeep, et al.
Published: (2026)
Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism
by: Cui, Chenwei, et al.
Published: (2026)
by: Cui, Chenwei, et al.
Published: (2026)
PRUE: A Practical Recipe for Field Boundary Segmentation at Scale
by: Muhawenayo, Gedeon, et al.
Published: (2026)
by: Muhawenayo, Gedeon, et al.
Published: (2026)
The first global agricultural field boundary map at 10m resolution
by: Robinson, Caleb, et al.
Published: (2026)
by: Robinson, Caleb, et al.
Published: (2026)
HyperMLP: An Integrated Perspective for Sequence Modeling
by: Lu, Jiecheng, et al.
Published: (2026)
by: Lu, Jiecheng, et al.
Published: (2026)
Incorporating Exponential Smoothing into MLP: A Simple but Effective Sequence Model
by: Chu, Jiqun, et al.
Published: (2024)
by: Chu, Jiqun, et al.
Published: (2024)
Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models
by: Badger, Benjamin L., et al.
Published: (2026)
by: Badger, Benjamin L., et al.
Published: (2026)
Understanding MLP-Mixer as a Wide and Sparse MLP
by: Hayase, Tomohiro, et al.
Published: (2023)
by: Hayase, Tomohiro, et al.
Published: (2023)
PreMixer: MLP-Based Pre-training Enhanced MLP-Mixers for Large-scale Traffic Forecasting
by: Zhang, Tongtong, et al.
Published: (2024)
by: Zhang, Tongtong, et al.
Published: (2024)
GraphMLP: A Graph MLP-Like Architecture for 3D Human Pose Estimation
by: Li, Wenhao, et al.
Published: (2022)
by: Li, Wenhao, et al.
Published: (2022)
Multi-Region Transfer Learning for Segmentation of Crop Field Boundaries in Satellite Images with Limited Labels
by: Kerner, Hannah, et al.
Published: (2024)
by: Kerner, Hannah, et al.
Published: (2024)
DPA: A one-stop metric to measure bias amplification in classification datasets
by: Tokas, Bhanu, et al.
Published: (2024)
by: Tokas, Bhanu, et al.
Published: (2024)
Alpha Excel Benchmark
by: Noever, David, et al.
Published: (2025)
by: Noever, David, et al.
Published: (2025)
Language Models "Grok" to Copy
by: Lv, Ang, et al.
Published: (2024)
by: Lv, Ang, et al.
Published: (2024)
ExcelFormer: A neural network surpassing GBDTs on tabular data
by: Chen, Jintai, et al.
Published: (2023)
by: Chen, Jintai, et al.
Published: (2023)
TimeDistill: Efficient Long-Term Time Series Forecasting with MLP via Cross-Architecture Distillation
by: Ni, Juntong, et al.
Published: (2025)
by: Ni, Juntong, et al.
Published: (2025)
From MLP to NeoMLP: Leveraging Self-Attention for Neural Fields
by: Kofinas, Miltiadis, et al.
Published: (2024)
by: Kofinas, Miltiadis, et al.
Published: (2024)
Approximation Rate of the Transformer Architecture for Sequence Modeling
by: Jiang, Haotian, et al.
Published: (2023)
by: Jiang, Haotian, et al.
Published: (2023)
Evolution Meets Diffusion: Efficient Neural Architecture Generation
by: Zhou, Bingye, et al.
Published: (2025)
by: Zhou, Bingye, et al.
Published: (2025)
A Simple State Space Model Excels at Multivariate Time Series Classification
by: Saadatmand, Hassan, et al.
Published: (2026)
by: Saadatmand, Hassan, et al.
Published: (2026)
Domain-Specific Pretraining of Language Models: A Comparative Study in the Medical Field
by: Kerner, Tobias
Published: (2024)
by: Kerner, Tobias
Published: (2024)
Architectural and Inferential Inductive Biases For Exchangeable Sequence Modeling
by: Mittal, Daksh, et al.
Published: (2025)
by: Mittal, Daksh, et al.
Published: (2025)
SE-MLP Model for Predicting Prior Acceleration Features in Penetration Signals
by: Li, Yankang, et al.
Published: (2025)
by: Li, Yankang, et al.
Published: (2025)
Tiny Transformers Excel at Sentence Compression
by: Belcak, Peter, et al.
Published: (2024)
by: Belcak, Peter, et al.
Published: (2024)
Rethinking the shape convention of an MLP
by: Chen, Meng-Hsi, et al.
Published: (2025)
by: Chen, Meng-Hsi, et al.
Published: (2025)
Spatial Transfer Learning with Simple MLP
by: Yang, Hongjian
Published: (2024)
by: Yang, Hongjian
Published: (2024)
Hybrid(Penalized Regression and MLP) Models for Outcome Prediction in HDLSS Health Data
by: K, Mithra D
Published: (2025)
by: K, Mithra D
Published: (2025)
DICE: Diffusion Large Language Models Excel at Generating CUDA Kernels
by: Bai, Haolei, et al.
Published: (2026)
by: Bai, Haolei, et al.
Published: (2026)
Bridging KAN and MLP: MJKAN, a Hybrid Architecture with Both Efficiency and Expressiveness
by: Joo, Hanseon, et al.
Published: (2025)
by: Joo, Hanseon, et al.
Published: (2025)
Maintaining and Managing Road Quality:Using MLP and DNN
by: Maotwana, Makgotso Jacqueline
Published: (2024)
by: Maotwana, Makgotso Jacqueline
Published: (2024)
Temporal Graph MLP Mixer for Spatio-Temporal Forecasting
by: Bilal, Muhammad, et al.
Published: (2025)
by: Bilal, Muhammad, et al.
Published: (2025)
Can an MLP Absorb Its Own Skip Connection?
by: Mijoski, Antonij, et al.
Published: (2026)
by: Mijoski, Antonij, et al.
Published: (2026)
Efficient LLMs with AMP: Attention Heads and MLP Pruning
by: Mugnaini, Leandro Giusti, et al.
Published: (2025)
by: Mugnaini, Leandro Giusti, et al.
Published: (2025)
End to End Autoencoder MLP Framework for Sepsis Prediction
by: Cai, Hejiang, et al.
Published: (2025)
by: Cai, Hejiang, et al.
Published: (2025)
Sequence-to-Sequence Models with Attention Mechanistically Map to the Architecture of Human Memory Search
by: Salvatore, Nikolaus, et al.
Published: (2025)
by: Salvatore, Nikolaus, et al.
Published: (2025)
Avocado Price Prediction Using a Hybrid Deep Learning Model: TCN-MLP-Attention Architecture
by: Zhang, Linwei, et al.
Published: (2025)
by: Zhang, Linwei, et al.
Published: (2025)
Teaching MLP More Graph Information: A Three-stage Multitask Knowledge Distillation Framework
by: Li, Junxian, et al.
Published: (2024)
by: Li, Junxian, et al.
Published: (2024)
PowerMLP: An Efficient Version of KAN
by: Qiu, Ruichen, et al.
Published: (2024)
by: Qiu, Ruichen, et al.
Published: (2024)
KAN or MLP: A Fairer Comparison
by: Yu, Runpeng, et al.
Published: (2024)
by: Yu, Runpeng, et al.
Published: (2024)
Similar Items
-
How Does the Spatial Distribution of Pre-training Data Affect Geospatial Foundation Models?
by: Purohit, Mirali, et al.
Published: (2025) -
Pretrain Where? Investigating How Pretraining Data Diversity Impacts Geospatial Foundation Model Performance
by: Kaur, Amandeep, et al.
Published: (2026) -
Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism
by: Cui, Chenwei, et al.
Published: (2026) -
PRUE: A Practical Recipe for Field Boundary Segmentation at Scale
by: Muhawenayo, Gedeon, et al.
Published: (2026) -
The first global agricultural field boundary map at 10m resolution
by: Robinson, Caleb, et al.
Published: (2026)