Multi-Token Prediction via Self-Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kirchenbauer, John, Hans, Abhimanyu, Bartoldson, Brian, Goldblum, Micah, Panda, Ashwinee, Goldstein, Tom
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913058141503488
author Kirchenbauer, John
Hans, Abhimanyu
Bartoldson, Brian
Goldblum, Micah
Panda, Ashwinee
Goldstein, Tom
author_facet Kirchenbauer, John
Hans, Abhimanyu
Bartoldson, Brian
Goldblum, Micah
Panda, Ashwinee
Goldstein, Tom
contents Existing techniques for accelerating language model inference, such as speculative decoding, require training auxiliary speculator models and building and deploying complex inference pipelines. We consider a new approach for converting a pretrained autoregressive language model from a slow single next token prediction model into a fast standalone multi-token prediction model using a simple online distillation objective. The final model retains the exact same implementation as the pretrained initial checkpoint and is deployable without the addition of any auxiliary verifier or other specialized inference code. Our method produces models that decode more than $3\times$ faster at $<5\%$ drop in accuracy on GSM8K relative to the single token decoding performance of the same checkpoint.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06019
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multi-Token Prediction via Self-Distillation
Kirchenbauer, John
Hans, Abhimanyu
Bartoldson, Brian
Goldblum, Micah
Panda, Ashwinee
Goldstein, Tom
Computation and Language
Machine Learning
Existing techniques for accelerating language model inference, such as speculative decoding, require training auxiliary speculator models and building and deploying complex inference pipelines. We consider a new approach for converting a pretrained autoregressive language model from a slow single next token prediction model into a fast standalone multi-token prediction model using a simple online distillation objective. The final model retains the exact same implementation as the pretrained initial checkpoint and is deployable without the addition of any auxiliary verifier or other specialized inference code. Our method produces models that decode more than $3\times$ faster at $<5\%$ drop in accuracy on GSM8K relative to the single token decoding performance of the same checkpoint.
title Multi-Token Prediction via Self-Distillation
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2602.06019