Exploring the Hidden Capacity of LLMs for One-Step Text Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mezentsev, Gleb, Oseledets, Ivan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914128976674816
author Mezentsev, Gleb
Oseledets, Ivan
author_facet Mezentsev, Gleb
Oseledets, Ivan
contents A recent study showed that large language models (LLMs) can reconstruct surprisingly long texts - up to thousands of tokens - via autoregressive generation from just one trained input embedding. In this work, we explore whether autoregressive decoding is essential for such reconstruction. We show that frozen LLMs can generate hundreds of accurate tokens in just one token-parallel forward pass, when provided with only two learned embeddings. This reveals a surprising and underexplored multi-token generation capability of autoregressive LLMs. We examine these embeddings and characterize the information they encode. We also empirically show that, although these representations are not unique for a given text, they form connected and local regions in embedding space - suggesting the potential to train a practical encoder. The existence of such representations hints that multi-token generation may be natively accessible in off-the-shelf LLMs via a learned input encoder, eliminating heavy retraining and helping to overcome the fundamental bottleneck of autoregressive decoding while reusing already-trained models.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21189
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring the Hidden Capacity of LLMs for One-Step Text Generation
Mezentsev, Gleb
Oseledets, Ivan
Computation and Language
Artificial Intelligence
Machine Learning
A recent study showed that large language models (LLMs) can reconstruct surprisingly long texts - up to thousands of tokens - via autoregressive generation from just one trained input embedding. In this work, we explore whether autoregressive decoding is essential for such reconstruction. We show that frozen LLMs can generate hundreds of accurate tokens in just one token-parallel forward pass, when provided with only two learned embeddings. This reveals a surprising and underexplored multi-token generation capability of autoregressive LLMs. We examine these embeddings and characterize the information they encode. We also empirically show that, although these representations are not unique for a given text, they form connected and local regions in embedding space - suggesting the potential to train a practical encoder. The existence of such representations hints that multi-token generation may be natively accessible in off-the-shelf LLMs via a learned input encoder, eliminating heavy retraining and helping to overcome the fundamental bottleneck of autoregressive decoding while reusing already-trained models.
title Exploring the Hidden Capacity of LLMs for One-Step Text Generation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.21189