STEM: Scaling Transformers with Embedding Modules

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sadhukhan, Ranajoy, Cao, Sheng, Dong, Harry, Zhao, Changsheng, Purpura-Pontoniere, Attiano, Tian, Yuandong, Liu, Zechun, Chen, Beidi
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912826516307968
author Sadhukhan, Ranajoy
Cao, Sheng
Dong, Harry
Zhao, Changsheng
Purpura-Pontoniere, Attiano
Tian, Yuandong
Liu, Zechun
Chen, Beidi
author_facet Sadhukhan, Ranajoy
Cao, Sheng
Dong, Harry
Zhao, Changsheng
Purpura-Pontoniere, Attiano
Tian, Yuandong
Liu, Zechun
Chen, Beidi
contents Fine-grained sparsity promises higher parametric capacity without proportional per-token compute, but often suffers from training instability, load balancing, and communication overhead. We introduce STEM (Scaling Transformers with Embedding Modules), a static, token-indexed approach that replaces the FFN up-projection with a layer-local embedding lookup while keeping the gate and down-projection dense. This removes runtime routing, enables CPU offload with asynchronous prefetch, and decouples capacity from both per-token FLOPs and cross-device communication. Empirically, STEM trains stably despite extreme sparsity. It improves downstream performance over dense baselines while reducing per-token FLOPs and parameter accesses (eliminating roughly one-third of FFN parameters). STEM learns embedding spaces with large angular spread which enhances its knowledge storage capacity. More interestingly, this enhanced knowledge capacity comes with better interpretability. The token-indexed nature of STEM embeddings allows simple ways to perform knowledge editing and knowledge injection in an interpretable manner without any intervention in the input text or additional computation. In addition, STEM strengthens long-context performance: as sequence length grows, more distinct parameters are activated, yielding practical test-time capacity scaling. Across 350M and 1B model scales, STEM delivers up to ~3--4% accuracy improvements overall, with notable gains on knowledge and reasoning-heavy benchmarks (ARC-Challenge, OpenBookQA, GSM8K, MMLU). Overall, STEM is an effective way of scaling parametric memory while providing better interpretability, better training stability and improved efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2601_10639
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle STEM: Scaling Transformers with Embedding Modules
Sadhukhan, Ranajoy
Cao, Sheng
Dong, Harry
Zhao, Changsheng
Purpura-Pontoniere, Attiano
Tian, Yuandong
Liu, Zechun
Chen, Beidi
Machine Learning
Fine-grained sparsity promises higher parametric capacity without proportional per-token compute, but often suffers from training instability, load balancing, and communication overhead. We introduce STEM (Scaling Transformers with Embedding Modules), a static, token-indexed approach that replaces the FFN up-projection with a layer-local embedding lookup while keeping the gate and down-projection dense. This removes runtime routing, enables CPU offload with asynchronous prefetch, and decouples capacity from both per-token FLOPs and cross-device communication. Empirically, STEM trains stably despite extreme sparsity. It improves downstream performance over dense baselines while reducing per-token FLOPs and parameter accesses (eliminating roughly one-third of FFN parameters). STEM learns embedding spaces with large angular spread which enhances its knowledge storage capacity. More interestingly, this enhanced knowledge capacity comes with better interpretability. The token-indexed nature of STEM embeddings allows simple ways to perform knowledge editing and knowledge injection in an interpretable manner without any intervention in the input text or additional computation. In addition, STEM strengthens long-context performance: as sequence length grows, more distinct parameters are activated, yielding practical test-time capacity scaling. Across 350M and 1B model scales, STEM delivers up to ~3--4% accuracy improvements overall, with notable gains on knowledge and reasoning-heavy benchmarks (ARC-Challenge, OpenBookQA, GSM8K, MMLU). Overall, STEM is an effective way of scaling parametric memory while providing better interpretability, better training stability and improved efficiency.
title STEM: Scaling Transformers with Embedding Modules
topic Machine Learning
url https://arxiv.org/abs/2601.10639