SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Csordás, Róbert, Piękos, Piotr, Irie, Kazuki, Schmidhuber, Jürgen |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoEUT: Mixture-of-Experts Universal Transformers
by: Csordás, Róbert, et al.
Published: (2024)
by: Csordás, Róbert, et al.
Published: (2024)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
by: Piękos, Piotr, et al.
Published: (2025)
by: Piękos, Piotr, et al.
Published: (2025)
Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space
by: Liu, Houjun, et al.
Published: (2025)
by: Liu, Houjun, et al.
Published: (2025)
Who invented deep residual learning?
by: Schmidhuber, Juergen
Published: (2025)
by: Schmidhuber, Juergen
Published: (2025)
Learning to Forget: Continual Learning with Adaptive Weight Decay
by: Ramesh, Aditya A., et al.
Published: (2026)
by: Ramesh, Aditya A., et al.
Published: (2026)
Parameter-Efficient Fine-Tuning of LLMs with Mixture of Space Experts
by: Zhang, Buze, et al.
Published: (2026)
by: Zhang, Buze, et al.
Published: (2026)
Recurrent Complex-Weighted Autoencoders for Unsupervised Object Discovery
by: Gopalakrishnan, Anand, et al.
Published: (2024)
by: Gopalakrishnan, Anand, et al.
Published: (2024)
GLU Attention Improve Transformer
by: Wang, Zehao
Published: (2025)
by: Wang, Zehao
Published: (2025)
Mixture of Experts with Soft Nearest Neighbor Loss: Resolving Expert Collapse via Representation Disentanglement
by: Agarap, Abien Fred, et al.
Published: (2026)
by: Agarap, Abien Fred, et al.
Published: (2026)
Do Language Models Use Their Depth Efficiently?
by: Csordás, Róbert, et al.
Published: (2025)
by: Csordás, Róbert, et al.
Published: (2025)
Affinity Is Not Enough: Recovering the Free Energy Principle in Mixture-of-Experts
by: Wong, Man Yung
Published: (2026)
by: Wong, Man Yung
Published: (2026)
A Gated Residual Kolmogorov-Arnold Networks for Mixtures of Experts
by: Inzirillo, Hugo, et al.
Published: (2024)
by: Inzirillo, Hugo, et al.
Published: (2024)
Decomposing Evolutionary Mixture-of-LoRA Architectures: The Routing Lever, the Lifecycle Penalty, and a Substrate-Conditional Boundary
by: Kumaresan, Ramchand
Published: (2026)
by: Kumaresan, Ramchand
Published: (2026)
NOBLE: Accelerating Transformers with Nonlinear Low-Rank Branches
by: Smith, Ethan
Published: (2026)
by: Smith, Ethan
Published: (2026)
On the Power of Convolution Augmented Transformer
by: Li, Mingchen, et al.
Published: (2024)
by: Li, Mingchen, et al.
Published: (2024)
BrainTransformers: SNN-LLM
by: Tang, Zhengzheng, et al.
Published: (2024)
by: Tang, Zhengzheng, et al.
Published: (2024)
HAT: Hardware-Aware Transformers for Efficient Natural Language Processing
by: Wang, Hanrui, et al.
Published: (2020)
by: Wang, Hanrui, et al.
Published: (2020)
A Hormone-inspired Emotion Layer for Transformer language models (HELT)
by: Reda, Eslam, et al.
Published: (2026)
by: Reda, Eslam, et al.
Published: (2026)
Sorbet: A Neuromorphic Hardware-Compatible Transformer-Based Spiking Language Model
by: Tang, Kaiwen, et al.
Published: (2024)
by: Tang, Kaiwen, et al.
Published: (2024)
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
by: Csordás, Róbert, et al.
Published: (2024)
by: Csordás, Róbert, et al.
Published: (2024)
Accelerating Vehicle Routing via AI-Initialized Genetic Algorithms
by: Greenberg, Ido, et al.
Published: (2025)
by: Greenberg, Ido, et al.
Published: (2025)
Selective Synchronization Attention
by: Hays, Hasi
Published: (2026)
by: Hays, Hasi
Published: (2026)
Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers
by: Xu, Boxun, et al.
Published: (2024)
by: Xu, Boxun, et al.
Published: (2024)
SpikeGraphormer: A High-Performance Graph Transformer with Spiking Graph Attention
by: Sun, Yundong, et al.
Published: (2024)
by: Sun, Yundong, et al.
Published: (2024)
Hierarchical Kernel Transformer: Multi-Scale Attention with an Information-Theoretic Approximation Analysis
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
Accelerating Training Speed of Tiny Recursive Models with Curriculum Guided Adaptive Recursion
by: Qasim, Kaleem Ullah, et al.
Published: (2025)
by: Qasim, Kaleem Ullah, et al.
Published: (2025)
Large Language Models Suffer From Their Own Output: An Analysis of the Self-Consuming Training Loop
by: Briesch, Martin, et al.
Published: (2023)
by: Briesch, Martin, et al.
Published: (2023)
ComplicaCode: Enhancing Disease Complication Detection in Electronic Health Records through ICD Path Generation
by: Zhou, Xiaofan
Published: (2023)
by: Zhou, Xiaofan
Published: (2023)
Improving Language Plasticity via Pretraining with Active Forgetting
by: Chen, Yihong, et al.
Published: (2023)
by: Chen, Yihong, et al.
Published: (2023)
SpikeGPT: Generative Pre-trained Language Model with Spiking Neural Networks
by: Zhu, Rui-Jie, et al.
Published: (2023)
by: Zhu, Rui-Jie, et al.
Published: (2023)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
by: Yu, Bohan, et al.
Published: (2025)
by: Yu, Bohan, et al.
Published: (2025)
An In-depth Walkthrough on Evolution of Neural Machine Translation
by: Jagtap, Rohan, et al.
Published: (2020)
by: Jagtap, Rohan, et al.
Published: (2020)
Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
by: Dong, Peijie, et al.
Published: (2024)
by: Dong, Peijie, et al.
Published: (2024)
SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking Mechanisms
by: Xing, Xingrun, et al.
Published: (2024)
by: Xing, Xingrun, et al.
Published: (2024)
Hysteresis Activation Function for Efficient Inference
by: Kimhi, Moshe, et al.
Published: (2024)
by: Kimhi, Moshe, et al.
Published: (2024)
Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models
by: Majumdar, Somshubra, et al.
Published: (2024)
by: Majumdar, Somshubra, et al.
Published: (2024)
Pre-trained Language Models Learn Remarkably Accurate Representations of Numbers
by: Kadlčík, Marek, et al.
Published: (2025)
by: Kadlčík, Marek, et al.
Published: (2025)
Large Language Models for Tuning Evolution Strategies
by: Kramer, Oliver
Published: (2024)
by: Kramer, Oliver
Published: (2024)
AP-BMM: Approximating Capability-Cost Pareto Sets of LLMs via Asynchronous Prior-Guided Bayesian Model Merging
by: Chen, Kesheng, et al.
Published: (2025)
by: Chen, Kesheng, et al.
Published: (2025)
SpikingSSMs: Learning Long Sequences with Sparse and Parallel Spiking State Space Models
by: Shen, Shuaijie, et al.
Published: (2024)
by: Shen, Shuaijie, et al.
Published: (2024)
Similar Items
-
MoEUT: Mixture-of-Experts Universal Transformers
by: Csordás, Róbert, et al.
Published: (2024) -
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
by: Piękos, Piotr, et al.
Published: (2025) -
Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space
by: Liu, Houjun, et al.
Published: (2025) -
Who invented deep residual learning?
by: Schmidhuber, Juergen
Published: (2025) -
Learning to Forget: Continual Learning with Adaptive Weight Decay
by: Ramesh, Aditya A., et al.
Published: (2026)