Saved in:
Bibliographic Details
Main Authors: Yang, Yebin, Wu, Huaijin, Guo, Fu, Yao, Lin, Qin, Xiaohan, Wang, Jingzhi, Zhang, Debing, Yan, Junchi
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.00800
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910007514103808
author Yang, Yebin
Wu, Huaijin
Guo, Fu
Yao, Lin
Qin, Xiaohan
Wang, Jingzhi
Zhang, Debing
Yan, Junchi
author_facet Yang, Yebin
Wu, Huaijin
Guo, Fu
Yao, Lin
Qin, Xiaohan
Wang, Jingzhi
Zhang, Debing
Yan, Junchi
contents LLMs have traditionally scaled along dense dimensions, where performance is coupled with near-linear increases in computational cost. While MoE decouples capacity from compute, it introduces large memory overhead and hardware efficiency challenges. To overcome these, we propose token-indexed parameters as a novel, orthogonal scaling axis that decouple model capacity from FLOPs. Specifically, we introduce Joint-Token (JTok) and Mixture of Joint-Token (JTok-M), which augment Transformer layers with modulation vectors retrieved from auxiliary embedding tables. These vectors modulate the backbone via lightweight, element-wise operations, incurring negligible FLOPs overhead. Extensive experiments on both dense and MoE backbones, spanning from 650M (190M + 460M embedding) to 61B (17B + 44B embedding) total parameters, demonstrate that our approach consistently reduces validation loss and significantly improves downstream task performance (e.g., +4.1 on MMLU, +8.3 on ARC, +8.9 on CEval). Rigorous isoFLOPs analysis further confirms that JTok-M fundamentally shifts the quality-compute Pareto frontier, achieving comparable model quality with 35% less compute relative to vanilla MoE architectures, and we validate that token-indexed parameters exhibit a predictable power-law scaling behavior. Moreover, our efficient implementation ensures that the overhead introduced by JTok and JTok-M remains marginal.
format Preprint
id arxiv_https___arxiv_org_abs_2602_00800
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle JTok: On Token Embedding as another Axis of Scaling Law via Joint Token Self-modulation
Yang, Yebin
Wu, Huaijin
Guo, Fu
Yao, Lin
Qin, Xiaohan
Wang, Jingzhi
Zhang, Debing
Yan, Junchi
Machine Learning
Artificial Intelligence
LLMs have traditionally scaled along dense dimensions, where performance is coupled with near-linear increases in computational cost. While MoE decouples capacity from compute, it introduces large memory overhead and hardware efficiency challenges. To overcome these, we propose token-indexed parameters as a novel, orthogonal scaling axis that decouple model capacity from FLOPs. Specifically, we introduce Joint-Token (JTok) and Mixture of Joint-Token (JTok-M), which augment Transformer layers with modulation vectors retrieved from auxiliary embedding tables. These vectors modulate the backbone via lightweight, element-wise operations, incurring negligible FLOPs overhead. Extensive experiments on both dense and MoE backbones, spanning from 650M (190M + 460M embedding) to 61B (17B + 44B embedding) total parameters, demonstrate that our approach consistently reduces validation loss and significantly improves downstream task performance (e.g., +4.1 on MMLU, +8.3 on ARC, +8.9 on CEval). Rigorous isoFLOPs analysis further confirms that JTok-M fundamentally shifts the quality-compute Pareto frontier, achieving comparable model quality with 35% less compute relative to vanilla MoE architectures, and we validate that token-indexed parameters exhibit a predictable power-law scaling behavior. Moreover, our efficient implementation ensures that the overhead introduced by JTok and JTok-M remains marginal.
title JTok: On Token Embedding as another Axis of Scaling Law via Joint Token Self-modulation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.00800