Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.00800 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910007514103808 |
|---|---|
| author | Yang, Yebin Wu, Huaijin Guo, Fu Yao, Lin Qin, Xiaohan Wang, Jingzhi Zhang, Debing Yan, Junchi |
| author_facet | Yang, Yebin Wu, Huaijin Guo, Fu Yao, Lin Qin, Xiaohan Wang, Jingzhi Zhang, Debing Yan, Junchi |
| contents | LLMs have traditionally scaled along dense dimensions, where performance is coupled with near-linear increases in computational cost. While MoE decouples capacity from compute, it introduces large memory overhead and hardware efficiency challenges. To overcome these, we propose token-indexed parameters as a novel, orthogonal scaling axis that decouple model capacity from FLOPs. Specifically, we introduce Joint-Token (JTok) and Mixture of Joint-Token (JTok-M), which augment Transformer layers with modulation vectors retrieved from auxiliary embedding tables. These vectors modulate the backbone via lightweight, element-wise operations, incurring negligible FLOPs overhead. Extensive experiments on both dense and MoE backbones, spanning from 650M (190M + 460M embedding) to 61B (17B + 44B embedding) total parameters, demonstrate that our approach consistently reduces validation loss and significantly improves downstream task performance (e.g., +4.1 on MMLU, +8.3 on ARC, +8.9 on CEval). Rigorous isoFLOPs analysis further confirms that JTok-M fundamentally shifts the quality-compute Pareto frontier, achieving comparable model quality with 35% less compute relative to vanilla MoE architectures, and we validate that token-indexed parameters exhibit a predictable power-law scaling behavior. Moreover, our efficient implementation ensures that the overhead introduced by JTok and JTok-M remains marginal. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_00800 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | JTok: On Token Embedding as another Axis of Scaling Law via Joint Token Self-modulation Yang, Yebin Wu, Huaijin Guo, Fu Yao, Lin Qin, Xiaohan Wang, Jingzhi Zhang, Debing Yan, Junchi Machine Learning Artificial Intelligence LLMs have traditionally scaled along dense dimensions, where performance is coupled with near-linear increases in computational cost. While MoE decouples capacity from compute, it introduces large memory overhead and hardware efficiency challenges. To overcome these, we propose token-indexed parameters as a novel, orthogonal scaling axis that decouple model capacity from FLOPs. Specifically, we introduce Joint-Token (JTok) and Mixture of Joint-Token (JTok-M), which augment Transformer layers with modulation vectors retrieved from auxiliary embedding tables. These vectors modulate the backbone via lightweight, element-wise operations, incurring negligible FLOPs overhead. Extensive experiments on both dense and MoE backbones, spanning from 650M (190M + 460M embedding) to 61B (17B + 44B embedding) total parameters, demonstrate that our approach consistently reduces validation loss and significantly improves downstream task performance (e.g., +4.1 on MMLU, +8.3 on ARC, +8.9 on CEval). Rigorous isoFLOPs analysis further confirms that JTok-M fundamentally shifts the quality-compute Pareto frontier, achieving comparable model quality with 35% less compute relative to vanilla MoE architectures, and we validate that token-indexed parameters exhibit a predictable power-law scaling behavior. Moreover, our efficient implementation ensures that the overhead introduced by JTok and JTok-M remains marginal. |
| title | JTok: On Token Embedding as another Axis of Scaling Law via Joint Token Self-modulation |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2602.00800 |