A Frustratingly Simple Decoding Method for Neural Text Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Haoran, Cai, Deng, Li, Huayang, Bi, Wei, Lam, Wai, Shi, Shuming
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929257186328576
author Yang, Haoran
Cai, Deng
Li, Huayang
Bi, Wei
Lam, Wai
Shi, Shuming
author_facet Yang, Haoran
Cai, Deng
Li, Huayang
Bi, Wei
Lam, Wai
Shi, Shuming
contents We introduce a frustratingly simple, super efficient and surprisingly effective decoding method, which we call Frustratingly Simple Decoding (FSD), for neural text generation. The idea behind FSD is straightforward: we build an anti-LM based on previously generated text and use this anti-LM to penalize future generation of what has been generated. The anti-LM can be implemented as simple as an n-gram language model or a vectorized variant. In this way, FSD introduces no extra model parameters and negligible computational overhead (FSD can be as fast as greedy search). Despite the simplicity, FSD is surprisingly effective; Experiments show that FSD can outperform the canonical methods to date (i.e., nucleus sampling) as well as several strong baselines that were proposed recently.
format Preprint
id arxiv_https___arxiv_org_abs_2305_12675
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A Frustratingly Simple Decoding Method for Neural Text Generation
Yang, Haoran
Cai, Deng
Li, Huayang
Bi, Wei
Lam, Wai
Shi, Shuming
Computation and Language
We introduce a frustratingly simple, super efficient and surprisingly effective decoding method, which we call Frustratingly Simple Decoding (FSD), for neural text generation. The idea behind FSD is straightforward: we build an anti-LM based on previously generated text and use this anti-LM to penalize future generation of what has been generated. The anti-LM can be implemented as simple as an n-gram language model or a vectorized variant. In this way, FSD introduces no extra model parameters and negligible computational overhead (FSD can be as fast as greedy search). Despite the simplicity, FSD is surprisingly effective; Experiments show that FSD can outperform the canonical methods to date (i.e., nucleus sampling) as well as several strong baselines that were proposed recently.
title A Frustratingly Simple Decoding Method for Neural Text Generation
topic Computation and Language
url https://arxiv.org/abs/2305.12675