MoEUT: Mixture-of-Experts Universal Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Csordás, Róbert, Irie, Kazuki, Schmidhuber, Jürgen, Potts, Christopher, Manning, Christopher D. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
by: Csordás, Róbert, et al.
Published: (2023)
by: Csordás, Róbert, et al.
Published: (2023)
Do Language Models Use Their Depth Efficiently?
by: Csordás, Róbert, et al.
Published: (2025)
by: Csordás, Róbert, et al.
Published: (2025)
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
by: Csordás, Róbert, et al.
Published: (2024)
by: Csordás, Róbert, et al.
Published: (2024)
Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space
by: Liu, Houjun, et al.
Published: (2025)
by: Liu, Houjun, et al.
Published: (2025)
Attending to Graph Transformers
by: Müller, Luis, et al.
Published: (2023)
by: Müller, Luis, et al.
Published: (2023)
Parameter-Efficient Fine-Tuning of LLMs with Mixture of Space Experts
by: Zhang, Buze, et al.
Published: (2026)
by: Zhang, Buze, et al.
Published: (2026)
Towards graph neural networks for provably solving convex optimization problems
by: Qian, Chendi, et al.
Published: (2025)
by: Qian, Chendi, et al.
Published: (2025)
EOE: Evolutionary Optimization of Experts for Training Language Models
by: Chen, Yingshi
Published: (2025)
by: Chen, Yingshi
Published: (2025)
Who invented deep residual learning?
by: Schmidhuber, Juergen
Published: (2025)
by: Schmidhuber, Juergen
Published: (2025)
JPC: Flexible Inference for Predictive Coding Networks in JAX
by: Innocenti, Francesco, et al.
Published: (2024)
by: Innocenti, Francesco, et al.
Published: (2024)
Learning to Forget: Continual Learning with Adaptive Weight Decay
by: Ramesh, Aditya A., et al.
Published: (2026)
by: Ramesh, Aditya A., et al.
Published: (2026)
Investigating Recurrent Transformers with Dynamic Halt
by: Chowdhury, Jishnu Ray, et al.
Published: (2024)
by: Chowdhury, Jishnu Ray, et al.
Published: (2024)
Structure Development in List-Sorting Transformers
by: Urdshals, Einar, et al.
Published: (2025)
by: Urdshals, Einar, et al.
Published: (2025)
Brain-inspired Computational Intelligence via Predictive Coding
by: Salvatori, Tommaso, et al.
Published: (2023)
by: Salvatori, Tommaso, et al.
Published: (2023)
Understanding Transformer Optimization via Gradient Heterogeneity
by: Tomihari, Akiyoshi, et al.
Published: (2025)
by: Tomihari, Akiyoshi, et al.
Published: (2025)
Spiking Point Transformer for Point Cloud Classification
by: Wu, Peixi, et al.
Published: (2025)
by: Wu, Peixi, et al.
Published: (2025)
SGHormer: An Energy-Saving Graph Transformer Driven by Spikes
by: Zhang, Huizhe, et al.
Published: (2024)
by: Zhang, Huizhe, et al.
Published: (2024)
General-Purpose In-Context Learning by Meta-Learning Transformers
by: Kirsch, Louis, et al.
Published: (2022)
by: Kirsch, Louis, et al.
Published: (2022)
QSViT: A Methodology for Quantizing Spiking Vision Transformers
by: Putra, Rachmad Vidya Wicaksana, et al.
Published: (2025)
by: Putra, Rachmad Vidya Wicaksana, et al.
Published: (2025)
Decision SpikeFormer: Spike-Driven Transformer for Decision Making
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain
by: Kosowski, Adrian, et al.
Published: (2025)
by: Kosowski, Adrian, et al.
Published: (2025)
Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics
by: Verma, Lucky
Published: (2026)
by: Verma, Lucky
Published: (2026)
Universal Neurons in GPT-2: Emergence, Persistence, and Functional Impact
by: Nandan, Advey, et al.
Published: (2025)
by: Nandan, Advey, et al.
Published: (2025)
Why "classic" Transformers are shallow and how to make them go deep
by: Yu, Yueyao, et al.
Published: (2023)
by: Yu, Yueyao, et al.
Published: (2023)
Sparsity Moves Computation: How FFN Architecture Reshapes Attention in Small Transformers
by: Smithline, Gabriel, et al.
Published: (2026)
by: Smithline, Gabriel, et al.
Published: (2026)
Sample-based Dynamic Hierarchical Transformer with Layer and Head Flexibility via Contextual Bandit
by: Meng, Fanfei, et al.
Published: (2023)
by: Meng, Fanfei, et al.
Published: (2023)
Towards Universal Offline Black-Box Optimization via Learning Language Model Embeddings
by: Tan, Rong-Xi, et al.
Published: (2025)
by: Tan, Rong-Xi, et al.
Published: (2025)
Enabling Robust In-Context Memory and Rapid Task Adaptation in Transformers with Hebbian and Gradient-Based Plasticity
by: Chaudhary, Siddharth
Published: (2025)
by: Chaudhary, Siddharth
Published: (2025)
Decoding Listeners Identity: Person Identification from EEG Signals Using a Lightweight Spiking Transformer
by: Lin, Zheyuan, et al.
Published: (2025)
by: Lin, Zheyuan, et al.
Published: (2025)
Position: Graph Learning Will Lose Relevance Due To Poor Benchmarks
by: Bechler-Speicher, Maya, et al.
Published: (2025)
by: Bechler-Speicher, Maya, et al.
Published: (2025)
Adversarially Robust Spiking Neural Networks Through Conversion
by: Özdenizci, Ozan, et al.
Published: (2023)
by: Özdenizci, Ozan, et al.
Published: (2023)
Predicting Deterioration in Mild Cognitive Impairment with Survival Transformers, Extreme Gradient Boosting and Cox Proportional Hazard Modelling
by: Musto, Henry, et al.
Published: (2024)
by: Musto, Henry, et al.
Published: (2024)
Provably Optimal Memory Capacity for Modern Hopfield Models: Transformer-Compatible Dense Associative Memories as Spherical Codes
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Recurrent Complex-Weighted Autoencoders for Unsupervised Object Discovery
by: Gopalakrishnan, Anand, et al.
Published: (2024)
by: Gopalakrishnan, Anand, et al.
Published: (2024)
Parallel Algorithms for Exact Enumeration of Deep Neural Network Activation Regions
by: Drammis, Sabrina, et al.
Published: (2024)
by: Drammis, Sabrina, et al.
Published: (2024)
Large Language Models As Evolution Strategies
by: Lange, Robert Tjarko, et al.
Published: (2024)
by: Lange, Robert Tjarko, et al.
Published: (2024)
Stein Variational Evolution Strategies
by: Braun, Cornelius V., et al.
Published: (2024)
by: Braun, Cornelius V., et al.
Published: (2024)
A Scalable Hybrid Training Approach for Recurrent Spiking Neural Networks
by: Baronig, Maximilian, et al.
Published: (2025)
by: Baronig, Maximilian, et al.
Published: (2025)
GraphBench: Next-generation graph learning benchmarking
by: Stoll, Timo, et al.
Published: (2025)
by: Stoll, Timo, et al.
Published: (2025)
Fractional-order spike-timing-dependent gradient descent for multi-layer spiking neural networks
by: Yang, Yi, et al.
Published: (2024)
by: Yang, Yi, et al.
Published: (2024)
Similar Items
-
SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
by: Csordás, Róbert, et al.
Published: (2023) -
Do Language Models Use Their Depth Efficiently?
by: Csordás, Róbert, et al.
Published: (2025) -
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
by: Csordás, Róbert, et al.
Published: (2024) -
Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space
by: Liu, Houjun, et al.
Published: (2025) -
Attending to Graph Transformers
by: Müller, Luis, et al.
Published: (2023)