The Discrete Charm of the MLP: Binary Routing of Continuous Signals in Transformer Feed-Forward Layers
Fuente:
arXiv
Salvato in:
| Autore principale: | Balogh, Peter |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Half the Nonlinearity Is Wasted: Measuring and Reallocating the Transformer's MLP Budget
di: Balogh, Peter
Pubblicazione: (2026)
di: Balogh, Peter
Pubblicazione: (2026)
On the Role of Transformer Feed-Forward Layers in Nonlinear In-Context Learning
di: Sun, Haoyuan, et al.
Pubblicazione: (2025)
di: Sun, Haoyuan, et al.
Pubblicazione: (2025)
Counting in Small Transformers: The Delicate Interplay between Attention and Feed-Forward Layers
di: Behrens, Freya, et al.
Pubblicazione: (2024)
di: Behrens, Freya, et al.
Pubblicazione: (2024)
Merging Feed-Forward Sublayers for Compressed Transformers
di: Verma, Neha, et al.
Pubblicazione: (2025)
di: Verma, Neha, et al.
Pubblicazione: (2025)
The Anxiety of Influence: Bloom Filters in Transformer Attention Heads
di: Balogh, Peter
Pubblicazione: (2026)
di: Balogh, Peter
Pubblicazione: (2026)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
di: Bozic, Vukasin, et al.
Pubblicazione: (2023)
di: Bozic, Vukasin, et al.
Pubblicazione: (2023)
HorNets: Learning from Discrete and Continuous Signals with Routing Neural Networks
di: Koloski, Boshko, et al.
Pubblicazione: (2025)
di: Koloski, Boshko, et al.
Pubblicazione: (2025)
T-MLP: Tailed Multi-Layer Perceptron for Level-of-Detail Signal Representation
di: Yang, Chuanxiang, et al.
Pubblicazione: (2025)
di: Yang, Chuanxiang, et al.
Pubblicazione: (2025)
PolyGLU: State-Conditional Activation Routing in Transformer Feed-Forward Networks
di: Medeiros, Daniel Nobrega
Pubblicazione: (2026)
di: Medeiros, Daniel Nobrega
Pubblicazione: (2026)
Supernodes and Halos: Loss-Critical Hubs in LLM Feed-Forward Layers
di: Cherilyn, Audrey, et al.
Pubblicazione: (2026)
di: Cherilyn, Audrey, et al.
Pubblicazione: (2026)
MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers
di: Jaiswal, Ajay, et al.
Pubblicazione: (2026)
di: Jaiswal, Ajay, et al.
Pubblicazione: (2026)
Feed-Forward Latent Domain Adaptation
di: Bohdal, Ondrej, et al.
Pubblicazione: (2022)
di: Bohdal, Ondrej, et al.
Pubblicazione: (2022)
Darkness Visible: Reading the Exception Handler of a Language Model
di: Balogh, Peter
Pubblicazione: (2026)
di: Balogh, Peter
Pubblicazione: (2026)
SE-MLP Model for Predicting Prior Acceleration Features in Penetration Signals
di: Li, Yankang, et al.
Pubblicazione: (2025)
di: Li, Yankang, et al.
Pubblicazione: (2025)
Understanding MLP-Mixer as a Wide and Sparse MLP
di: Hayase, Tomohiro, et al.
Pubblicazione: (2023)
di: Hayase, Tomohiro, et al.
Pubblicazione: (2023)
Stress Detection Using PPG Signal and Combined Deep CNN-MLP Network
di: Hasanpoor, Yasin, et al.
Pubblicazione: (2024)
di: Hasanpoor, Yasin, et al.
Pubblicazione: (2024)
Do Neurons Dream of Primitive Operators? Wake-Sleep Compression Rediscovers Schank's Event Semantics
di: Balogh, Peter
Pubblicazione: (2026)
di: Balogh, Peter
Pubblicazione: (2026)
Coherent Feed Forward Quantum Neural Network
di: Singh, Utkarsh, et al.
Pubblicazione: (2024)
di: Singh, Utkarsh, et al.
Pubblicazione: (2024)
Optimizing Dense Feed-Forward Neural Networks
di: Balderas, Luis, et al.
Pubblicazione: (2023)
di: Balderas, Luis, et al.
Pubblicazione: (2023)
NerVE: Nonlinear Eigenspectrum Dynamics in LLM Feed-Forward Networks
di: Jha, Nandan Kumar, et al.
Pubblicazione: (2026)
di: Jha, Nandan Kumar, et al.
Pubblicazione: (2026)
In-Context Learning of a Linear Transformer Block: Benefits of the MLP Component and One-Step GD Initialization
di: Zhang, Ruiqi, et al.
Pubblicazione: (2024)
di: Zhang, Ruiqi, et al.
Pubblicazione: (2024)
LSTM VS. Feed-Forward Autoencoders for Unsupervised Fault Detection in Hydraulic Pumps
di: Sánchez, P., et al.
Pubblicazione: (2026)
di: Sánchez, P., et al.
Pubblicazione: (2026)
RadixMLP -- Intra-batch Deduplication for Causal Transformers
di: Feil, Michael, et al.
Pubblicazione: (2026)
di: Feil, Michael, et al.
Pubblicazione: (2026)
Feed-Forward Optimization With Delayed Feedback for Neural Network Training
di: Flügel, Katharina, et al.
Pubblicazione: (2023)
di: Flügel, Katharina, et al.
Pubblicazione: (2023)
From MLP to NeoMLP: Leveraging Self-Attention for Neural Fields
di: Kofinas, Miltiadis, et al.
Pubblicazione: (2024)
di: Kofinas, Miltiadis, et al.
Pubblicazione: (2024)
Energy-Efficient Supervised Learning with a Binary Stochastic Forward-Forward Algorithm
di: Jaiswal, Risi, et al.
Pubblicazione: (2025)
di: Jaiswal, Risi, et al.
Pubblicazione: (2025)
A Contrastive Symmetric Forward-Forward Algorithm (SFFA) for Continual Learning Tasks
di: Terres-Escudero, Erik B., et al.
Pubblicazione: (2024)
di: Terres-Escudero, Erik B., et al.
Pubblicazione: (2024)
Forward-Learned Discrete Diffusion: Learning how to noise to denoise faster
di: Bartosh, Grigory, et al.
Pubblicazione: (2026)
di: Bartosh, Grigory, et al.
Pubblicazione: (2026)
A Quasilinear Algorithm for Computing Higher-Order Derivatives of Deep Feed-Forward Neural Networks
di: Chickering, Kyle R.
Pubblicazione: (2024)
di: Chickering, Kyle R.
Pubblicazione: (2024)
Flash Multi-Head Feed-Forward Network
di: Zhang, Minshen, et al.
Pubblicazione: (2025)
di: Zhang, Minshen, et al.
Pubblicazione: (2025)
Feed-Forward Neural Networks as a Mixed-Integer Program
di: Aftabi, Navid, et al.
Pubblicazione: (2024)
di: Aftabi, Navid, et al.
Pubblicazione: (2024)
Diversity-Guided MLP Reduction for Efficient Large Vision Transformers
di: Shen, Chengchao, et al.
Pubblicazione: (2025)
di: Shen, Chengchao, et al.
Pubblicazione: (2025)
Rethinking the shape convention of an MLP
di: Chen, Meng-Hsi, et al.
Pubblicazione: (2025)
di: Chen, Meng-Hsi, et al.
Pubblicazione: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
di: Neo, Clement, et al.
Pubblicazione: (2024)
di: Neo, Clement, et al.
Pubblicazione: (2024)
Spatial Transfer Learning with Simple MLP
di: Yang, Hongjian
Pubblicazione: (2024)
di: Yang, Hongjian
Pubblicazione: (2024)
FFAvatar: Few-Shot, Feed-Forward, and Generalizable Avatar Reconstruction
di: Nguyen, Thuan Hoang, et al.
Pubblicazione: (2026)
di: Nguyen, Thuan Hoang, et al.
Pubblicazione: (2026)
PreMixer: MLP-Based Pre-training Enhanced MLP-Mixers for Large-scale Traffic Forecasting
di: Zhang, Tongtong, et al.
Pubblicazione: (2024)
di: Zhang, Tongtong, et al.
Pubblicazione: (2024)
GANFusion: Feed-Forward Text-to-3D with Diffusion in GAN Space
di: Attaiki, Souhaib, et al.
Pubblicazione: (2024)
di: Attaiki, Souhaib, et al.
Pubblicazione: (2024)
Enhancing Fast Feed Forward Networks with Load Balancing and a Master Leaf Node
di: Charalampopoulos, Andreas, et al.
Pubblicazione: (2024)
di: Charalampopoulos, Andreas, et al.
Pubblicazione: (2024)
How not to Stitch Representations to Measure Similarity: Task Loss Matching versus Direct Matching
di: Balogh, András, et al.
Pubblicazione: (2024)
di: Balogh, András, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Half the Nonlinearity Is Wasted: Measuring and Reallocating the Transformer's MLP Budget
di: Balogh, Peter
Pubblicazione: (2026) -
On the Role of Transformer Feed-Forward Layers in Nonlinear In-Context Learning
di: Sun, Haoyuan, et al.
Pubblicazione: (2025) -
Counting in Small Transformers: The Delicate Interplay between Attention and Feed-Forward Layers
di: Behrens, Freya, et al.
Pubblicazione: (2024) -
Merging Feed-Forward Sublayers for Compressed Transformers
di: Verma, Neha, et al.
Pubblicazione: (2025) -
The Anxiety of Influence: Bloom Filters in Transformer Attention Heads
di: Balogh, Peter
Pubblicazione: (2026)