Text Generation Beyond Discrete Token Sampling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhuang, Yufan, Liu, Liyuan, Singh, Chandan, Shang, Jingbo, Gao, Jianfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917034837671936
author Zhuang, Yufan
Liu, Liyuan
Singh, Chandan
Shang, Jingbo
Gao, Jianfeng
author_facet Zhuang, Yufan
Liu, Liyuan
Singh, Chandan
Shang, Jingbo
Gao, Jianfeng
contents In standard autoregressive generation, an LLM predicts the next-token distribution, samples a discrete token, and then discards the distribution, passing only the sampled token as new input. To preserve this distribution's rich information, we propose Mixture of Inputs (MoI), a training-free method for autoregressive generation. After generating a token following the standard paradigm, we construct a new input that blends the generated discrete token with the previously discarded token distribution. Specifically, we employ a Bayesian estimation method that treats the token distribution as the prior, the sampled token as the observation, and replaces the conventional one-hot vector with the continuous posterior expectation as the new model input. MoI allows the model to maintain a richer internal representation throughout the generation process, resulting in improved text quality and reasoning capabilities. On mathematical reasoning, code generation, and PhD-level QA tasks, MoI consistently improves performance across multiple models including QwQ-32B, Nemotron-Super-49B, Gemma-3-27B, and DAPO-Qwen-32B, with no additional training and negligible computational overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14827
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Text Generation Beyond Discrete Token Sampling
Zhuang, Yufan
Liu, Liyuan
Singh, Chandan
Shang, Jingbo
Gao, Jianfeng
Computation and Language
Artificial Intelligence
In standard autoregressive generation, an LLM predicts the next-token distribution, samples a discrete token, and then discards the distribution, passing only the sampled token as new input. To preserve this distribution's rich information, we propose Mixture of Inputs (MoI), a training-free method for autoregressive generation. After generating a token following the standard paradigm, we construct a new input that blends the generated discrete token with the previously discarded token distribution. Specifically, we employ a Bayesian estimation method that treats the token distribution as the prior, the sampled token as the observation, and replaces the conventional one-hot vector with the continuous posterior expectation as the new model input. MoI allows the model to maintain a richer internal representation throughout the generation process, resulting in improved text quality and reasoning capabilities. On mathematical reasoning, code generation, and PhD-level QA tasks, MoI consistently improves performance across multiple models including QwQ-32B, Nemotron-Super-49B, Gemma-3-27B, and DAPO-Qwen-32B, with no additional training and negligible computational overhead.
title Text Generation Beyond Discrete Token Sampling
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.14827