Half the Nonlinearity Is Wasted: Measuring and Reallocating the Transformer's MLP Budget
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Balogh, Peter |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Anxiety of Influence: Bloom Filters in Transformer Attention Heads
von: Balogh, Peter
Veröffentlicht: (2026)
von: Balogh, Peter
Veröffentlicht: (2026)
Darkness Visible: Reading the Exception Handler of a Language Model
von: Balogh, Peter
Veröffentlicht: (2026)
von: Balogh, Peter
Veröffentlicht: (2026)
The Larger the Better? Improved LLM Code-Generation via Budget Reallocation
von: Hassid, Michael, et al.
Veröffentlicht: (2024)
von: Hassid, Michael, et al.
Veröffentlicht: (2024)
The Discrete Charm of the MLP: Binary Routing of Continuous Signals in Transformer Feed-Forward Layers
von: Balogh, Peter
Veröffentlicht: (2026)
von: Balogh, Peter
Veröffentlicht: (2026)
HS-STaR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget Reallocation
von: Xiong, Feng, et al.
Veröffentlicht: (2025)
von: Xiong, Feng, et al.
Veröffentlicht: (2025)
Do Neurons Dream of Primitive Operators? Wake-Sleep Compression Rediscovers Schank's Event Semantics
von: Balogh, Peter
Veröffentlicht: (2026)
von: Balogh, Peter
Veröffentlicht: (2026)
In-Context Learning of a Linear Transformer Block: Benefits of the MLP Component and One-Step GD Initialization
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2024)
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2024)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
von: Tian, Yuandong, et al.
Veröffentlicht: (2023)
von: Tian, Yuandong, et al.
Veröffentlicht: (2023)
MLP-KAN: Unifying Deep Representation and Function Learning
von: He, Yunhong, et al.
Veröffentlicht: (2024)
von: He, Yunhong, et al.
Veröffentlicht: (2024)
HyperMLP: An Integrated Perspective for Sequence Modeling
von: Lu, Jiecheng, et al.
Veröffentlicht: (2026)
von: Lu, Jiecheng, et al.
Veröffentlicht: (2026)
MLP Fusion: Towards Efficient Fine-tuning of Dense and Mixture-of-Experts Language Models
von: Ai, Mengting, et al.
Veröffentlicht: (2023)
von: Ai, Mengting, et al.
Veröffentlicht: (2023)
Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models
von: Badger, Benjamin L., et al.
Veröffentlicht: (2026)
von: Badger, Benjamin L., et al.
Veröffentlicht: (2026)
Incorporating Exponential Smoothing into MLP: A Simple but Effective Sequence Model
von: Chu, Jiqun, et al.
Veröffentlicht: (2024)
von: Chu, Jiqun, et al.
Veröffentlicht: (2024)
SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models
von: Zhu, Hourun, et al.
Veröffentlicht: (2025)
von: Zhu, Hourun, et al.
Veröffentlicht: (2025)
ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation
von: Mei, Zhiyu, et al.
Veröffentlicht: (2024)
von: Mei, Zhiyu, et al.
Veröffentlicht: (2024)
Cognitive Fatigue in Autoregressive Transformers: Formalization and Measurement
von: Marwah, Riju, et al.
Veröffentlicht: (2026)
von: Marwah, Riju, et al.
Veröffentlicht: (2026)
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
The Missing Half: Unveiling Training-time Implicit Safety Risks Beyond Deployment
von: Zhang, Zhexin, et al.
Veröffentlicht: (2026)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2026)
Drift-to-Action Controllers: Budgeted Interventions with Online Risk Certificates
von: Lamaakal, Ismail, et al.
Veröffentlicht: (2026)
von: Lamaakal, Ismail, et al.
Veröffentlicht: (2026)
State Stream Transformer (SST) V2: Parallel Training of Nonlinear Recurrence for Latent Space Reasoning
von: Aviss, Thea
Veröffentlicht: (2026)
von: Aviss, Thea
Veröffentlicht: (2026)
Tiny Transformers Excel at Sentence Compression
von: Belcak, Peter, et al.
Veröffentlicht: (2024)
von: Belcak, Peter, et al.
Veröffentlicht: (2024)
Federated Language Models Under Bandwidth Budgets: Distillation Rates and Conformal Coverage
von: Dubey, Prasanjit, et al.
Veröffentlicht: (2026)
von: Dubey, Prasanjit, et al.
Veröffentlicht: (2026)
HiP-LoRA: Budgeted Spectral Plasticity for Robust Low-Rank Adaptation
von: Chen, Lixian, et al.
Veröffentlicht: (2026)
von: Chen, Lixian, et al.
Veröffentlicht: (2026)
Capability-Guided Compression: Toward Interpretability-Aware Budget Allocation for Large Language Models
von: Gupta, Rishaank
Veröffentlicht: (2026)
von: Gupta, Rishaank
Veröffentlicht: (2026)
Position-Agnostic Pre-Projection for Transformer Attention: Nonlinear Feature Construction and Content Skip Before Q/K/V
von: Shinde, Chirag
Veröffentlicht: (2026)
von: Shinde, Chirag
Veröffentlicht: (2026)
BAGEN: Are LLM Agents Budget-Aware?
von: Lin, Yuxiang, et al.
Veröffentlicht: (2026)
von: Lin, Yuxiang, et al.
Veröffentlicht: (2026)
Token-Budget-Aware LLM Reasoning
von: Han, Tingxu, et al.
Veröffentlicht: (2024)
von: Han, Tingxu, et al.
Veröffentlicht: (2024)
LLMs on a Budget? Say HOLA
von: Siddiqui, Zohaib Hasan, et al.
Veröffentlicht: (2025)
von: Siddiqui, Zohaib Hasan, et al.
Veröffentlicht: (2025)
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
von: Mohtashami, Amirkeivan, et al.
Veröffentlicht: (2023)
von: Mohtashami, Amirkeivan, et al.
Veröffentlicht: (2023)
LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction
von: Zhou, Enshuai, et al.
Veröffentlicht: (2026)
von: Zhou, Enshuai, et al.
Veröffentlicht: (2026)
Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling
von: Wang, Xinglin, et al.
Veröffentlicht: (2026)
von: Wang, Xinglin, et al.
Veröffentlicht: (2026)
Language Adaptation on a Tight Academic Compute Budget: Tokenizer Swapping Works and Pure bfloat16 Is Enough
von: Dobler, Konstantin, et al.
Veröffentlicht: (2024)
von: Dobler, Konstantin, et al.
Veröffentlicht: (2024)
SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
NOBLE: Accelerating Transformers with Nonlinear Low-Rank Branches
von: Smith, Ethan
Veröffentlicht: (2026)
von: Smith, Ethan
Veröffentlicht: (2026)
Trainable Transformer in Transformer
von: Panigrahi, Abhishek, et al.
Veröffentlicht: (2023)
von: Panigrahi, Abhishek, et al.
Veröffentlicht: (2023)
Budget-aware Test-time Scaling via Discriminative Verification
von: Montgomery, Kyle, et al.
Veröffentlicht: (2025)
von: Montgomery, Kyle, et al.
Veröffentlicht: (2025)
Optimizing Anytime Reasoning via Budget Relative Policy Optimization
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
Post-Trained MoE Can Skip Half Experts via Self-Distillation
von: Lv, Xingtai, et al.
Veröffentlicht: (2026)
von: Lv, Xingtai, et al.
Veröffentlicht: (2026)
When Numbers Tell Half the Story: Human-Metric Alignment in Topic Model Evaluation
von: Prouteau, Thibault, et al.
Veröffentlicht: (2026)
von: Prouteau, Thibault, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
The Anxiety of Influence: Bloom Filters in Transformer Attention Heads
von: Balogh, Peter
Veröffentlicht: (2026) -
Darkness Visible: Reading the Exception Handler of a Language Model
von: Balogh, Peter
Veröffentlicht: (2026) -
The Larger the Better? Improved LLM Code-Generation via Budget Reallocation
von: Hassid, Michael, et al.
Veröffentlicht: (2024) -
The Discrete Charm of the MLP: Binary Routing of Continuous Signals in Transformer Feed-Forward Layers
von: Balogh, Peter
Veröffentlicht: (2026) -
HS-STaR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget Reallocation
von: Xiong, Feng, et al.
Veröffentlicht: (2025)