Predicting the Order of Upcoming Tokens Improves Language Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zuhri, Zayd M. K., Fuadi, Erland Hilman, Aji, Alham Fikri
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915798992289792
author Zuhri, Zayd M. K.
Fuadi, Erland Hilman
Aji, Alham Fikri
author_facet Zuhri, Zayd M. K.
Fuadi, Erland Hilman
Aji, Alham Fikri
contents Multi-token prediction (MTP) has been proposed as an auxiliary objective to improve next-token prediction (NTP) in language model training but shows inconsistent improvements, underperforming in standard NLP benchmarks. We found MTP's exact future token prediction to be too difficult as an auxiliary loss. Instead, we propose token order prediction (TOP), which trains models to order upcoming tokens by their proximity using a learning-to-rank loss. TOP requires only a single additional unembedding layer compared to MTP's multiple transformer layers. We pretrain models of 340M, 1.8B, and 7B parameters using NTP, MTP, DeepSeek MTP (DS-MTP) and TOP objectives. The results of nine standard NLP benchmarks show that TOP overall outperforms NTP, MTP, and DS-MTP even at scale. TOP models with continued training on math and code also perform better on 4 relevant benchmarks. On the synthetic star graph task, TOP enables pathfinding on graphs where NTP, MTP, and DS-MTP fail. Our code is available at https://github.com/zaydzuhri/token-order-prediction
format Preprint
id arxiv_https___arxiv_org_abs_2508_19228
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Predicting the Order of Upcoming Tokens Improves Language Modeling
Zuhri, Zayd M. K.
Fuadi, Erland Hilman
Aji, Alham Fikri
Machine Learning
Multi-token prediction (MTP) has been proposed as an auxiliary objective to improve next-token prediction (NTP) in language model training but shows inconsistent improvements, underperforming in standard NLP benchmarks. We found MTP's exact future token prediction to be too difficult as an auxiliary loss. Instead, we propose token order prediction (TOP), which trains models to order upcoming tokens by their proximity using a learning-to-rank loss. TOP requires only a single additional unembedding layer compared to MTP's multiple transformer layers. We pretrain models of 340M, 1.8B, and 7B parameters using NTP, MTP, DeepSeek MTP (DS-MTP) and TOP objectives. The results of nine standard NLP benchmarks show that TOP overall outperforms NTP, MTP, and DS-MTP even at scale. TOP models with continued training on math and code also perform better on 4 relevant benchmarks. On the synthetic star graph task, TOP enables pathfinding on graphs where NTP, MTP, and DS-MTP fail. Our code is available at https://github.com/zaydzuhri/token-order-prediction
title Predicting the Order of Upcoming Tokens Improves Language Modeling
topic Machine Learning
url https://arxiv.org/abs/2508.19228