Improving Next Tokens via Second-to-Last Predictions with Generate and Refine

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteur principal: Schneider, Johannes
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913690183270400
author Schneider, Johannes
author_facet Schneider, Johannes
contents Autoregressive language models like GPT aim to predict next tokens, while autoencoding models such as BERT are trained on tasks such as predicting masked tokens. We train a decoder-only architecture for predicting the second to last token for a sequence of tokens. Our approach yields higher computational training efficiency than BERT-style models by employing a structured deterministic approach to masking tokens. We use our model to improve the next token predictions of a standard GPT by combining both predictions in a ``generate-then-refine'' approach. We demonstrate on different variants of GPT-2 and different datasets that (not unexpectedly) second to last token predictions are much more accurate, i.e., more than 15\% higher accuracy than standard next token predictions. The ``generate-then-refine'' approach also demonstrates notable improvements in next-token predictions, yielding smaller yet consistent and significant gains.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15661
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Next Tokens via Second-to-Last Predictions with Generate and Refine
Schneider, Johannes
Computation and Language
Machine Learning
Autoregressive language models like GPT aim to predict next tokens, while autoencoding models such as BERT are trained on tasks such as predicting masked tokens. We train a decoder-only architecture for predicting the second to last token for a sequence of tokens. Our approach yields higher computational training efficiency than BERT-style models by employing a structured deterministic approach to masking tokens. We use our model to improve the next token predictions of a standard GPT by combining both predictions in a ``generate-then-refine'' approach. We demonstrate on different variants of GPT-2 and different datasets that (not unexpectedly) second to last token predictions are much more accurate, i.e., more than 15\% higher accuracy than standard next token predictions. The ``generate-then-refine'' approach also demonstrates notable improvements in next-token predictions, yielding smaller yet consistent and significant gains.
title Improving Next Tokens via Second-to-Last Predictions with Generate and Refine
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2411.15661