Multiple Token Divergence: Measuring and Steering In-Context Computation Density

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Herrmann, Vincent, Alcaide, Eric, Wand, Michael, Schmidhuber, Jürgen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912793075122176
author Herrmann, Vincent
Alcaide, Eric
Wand, Michael
Schmidhuber, Jürgen
author_facet Herrmann, Vincent
Alcaide, Eric
Wand, Michael
Schmidhuber, Jürgen
contents Measuring the in-context computational effort of language models is a key challenge, as metrics like next-token loss fail to capture reasoning complexity. Prior methods based on latent state compressibility can be invasive and unstable. We propose Multiple Token Divergence (MTD), a simple measure of computational effort defined as the KL divergence between a model's full output distribution and that of a shallow, auxiliary prediction head. MTD can be computed directly from pre-trained models with multiple prediction heads, requiring no additional training. Building on this, we introduce Divergence Steering, a novel decoding method to control the computational character of generated text. We empirically show that MTD is more effective than prior methods at distinguishing complex tasks from simple ones. On mathematical reasoning benchmarks, MTD correlates positively with problem difficulty. Lower MTD is associated with more accurate reasoning. MTD provides a practical, lightweight tool for analyzing and steering the computational dynamics of language models.
format Preprint
id arxiv_https___arxiv_org_abs_2512_22944
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multiple Token Divergence: Measuring and Steering In-Context Computation Density
Herrmann, Vincent
Alcaide, Eric
Wand, Michael
Schmidhuber, Jürgen
Machine Learning
I.2.6
Measuring the in-context computational effort of language models is a key challenge, as metrics like next-token loss fail to capture reasoning complexity. Prior methods based on latent state compressibility can be invasive and unstable. We propose Multiple Token Divergence (MTD), a simple measure of computational effort defined as the KL divergence between a model's full output distribution and that of a shallow, auxiliary prediction head. MTD can be computed directly from pre-trained models with multiple prediction heads, requiring no additional training. Building on this, we introduce Divergence Steering, a novel decoding method to control the computational character of generated text. We empirically show that MTD is more effective than prior methods at distinguishing complex tasks from simple ones. On mathematical reasoning benchmarks, MTD correlates positively with problem difficulty. Lower MTD is associated with more accurate reasoning. MTD provides a practical, lightweight tool for analyzing and steering the computational dynamics of language models.
title Multiple Token Divergence: Measuring and Steering In-Context Computation Density
topic Machine Learning
I.2.6
url https://arxiv.org/abs/2512.22944