Enhancing Latent Computation in Transformers with Latent Tokens

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Yuchang, Chen, Yanxi, Li, Yaliang, Ding, Bolin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913845775171584
author Sun, Yuchang
Chen, Yanxi
Li, Yaliang
Ding, Bolin
author_facet Sun, Yuchang
Chen, Yanxi
Li, Yaliang
Ding, Bolin
contents Augmenting large language models (LLMs) with auxiliary tokens has emerged as a promising strategy for enhancing model performance. In this work, we introduce a lightweight method termed latent tokens; these are dummy tokens that may be non-interpretable in natural language but steer the autoregressive decoding process of a Transformer-based LLM via the attention mechanism. The proposed latent tokens can be seamlessly integrated with a pre-trained Transformer, trained in a parameter-efficient manner, and applied flexibly at inference time, while adding minimal complexity overhead to the existing infrastructure of standard Transformers. We propose several hypotheses about the underlying mechanisms of latent tokens and design synthetic tasks accordingly to verify them. Numerical results confirm that the proposed method noticeably outperforms the baselines, particularly in the out-of-distribution generalization scenarios, highlighting its potential in improving the adaptability of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12629
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Latent Computation in Transformers with Latent Tokens
Sun, Yuchang
Chen, Yanxi
Li, Yaliang
Ding, Bolin
Machine Learning
Computation and Language
Augmenting large language models (LLMs) with auxiliary tokens has emerged as a promising strategy for enhancing model performance. In this work, we introduce a lightweight method termed latent tokens; these are dummy tokens that may be non-interpretable in natural language but steer the autoregressive decoding process of a Transformer-based LLM via the attention mechanism. The proposed latent tokens can be seamlessly integrated with a pre-trained Transformer, trained in a parameter-efficient manner, and applied flexibly at inference time, while adding minimal complexity overhead to the existing infrastructure of standard Transformers. We propose several hypotheses about the underlying mechanisms of latent tokens and design synthetic tasks accordingly to verify them. Numerical results confirm that the proposed method noticeably outperforms the baselines, particularly in the out-of-distribution generalization scenarios, highlighting its potential in improving the adaptability of LLMs.
title Enhancing Latent Computation in Transformers with Latent Tokens
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2505.12629