Corrector Sampling in Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gat, Itai, Shaul, Neta, Singer, Uriel, Lipman, Yaron
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912417421721600
author Gat, Itai
Shaul, Neta
Singer, Uriel
Lipman, Yaron
author_facet Gat, Itai
Shaul, Neta
Singer, Uriel
Lipman, Yaron
contents Autoregressive language models accumulate errors due to their fixed, irrevocable left-to-right token generation. To address this, we propose a new sampling method called Resample-Previous-Tokens (RPT). RPT mitigates error accumulation by iteratively revisiting and potentially replacing tokens in a window of previously generated text. This method can be integrated into existing autoregressive models, preserving their next-token-prediction quality and speed. Fine-tuning a pretrained 8B parameter model with RPT for only 100B resulted in ~10% relative improvements on reasoning and coding benchmarks compared to the standard sampling.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06215
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Corrector Sampling in Language Models
Gat, Itai
Shaul, Neta
Singer, Uriel
Lipman, Yaron
Machine Learning
Computation and Language
Autoregressive language models accumulate errors due to their fixed, irrevocable left-to-right token generation. To address this, we propose a new sampling method called Resample-Previous-Tokens (RPT). RPT mitigates error accumulation by iteratively revisiting and potentially replacing tokens in a window of previously generated text. This method can be integrated into existing autoregressive models, preserving their next-token-prediction quality and speed. Fine-tuning a pretrained 8B parameter model with RPT for only 100B resulted in ~10% relative improvements on reasoning and coding benchmarks compared to the standard sampling.
title Corrector Sampling in Language Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2506.06215