A Frustratingly Simple Decoding Method for Neural Text Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929257186328576 |
|---|---|
| author | Yang, Haoran Cai, Deng Li, Huayang Bi, Wei Lam, Wai Shi, Shuming |
| author_facet | Yang, Haoran Cai, Deng Li, Huayang Bi, Wei Lam, Wai Shi, Shuming |
| contents | We introduce a frustratingly simple, super efficient and surprisingly effective decoding method, which we call Frustratingly Simple Decoding (FSD), for neural text generation. The idea behind FSD is straightforward: we build an anti-LM based on previously generated text and use this anti-LM to penalize future generation of what has been generated. The anti-LM can be implemented as simple as an n-gram language model or a vectorized variant. In this way, FSD introduces no extra model parameters and negligible computational overhead (FSD can be as fast as greedy search). Despite the simplicity, FSD is surprisingly effective; Experiments show that FSD can outperform the canonical methods to date (i.e., nucleus sampling) as well as several strong baselines that were proposed recently. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2305_12675 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | A Frustratingly Simple Decoding Method for Neural Text Generation Yang, Haoran Cai, Deng Li, Huayang Bi, Wei Lam, Wai Shi, Shuming Computation and Language We introduce a frustratingly simple, super efficient and surprisingly effective decoding method, which we call Frustratingly Simple Decoding (FSD), for neural text generation. The idea behind FSD is straightforward: we build an anti-LM based on previously generated text and use this anti-LM to penalize future generation of what has been generated. The anti-LM can be implemented as simple as an n-gram language model or a vectorized variant. In this way, FSD introduces no extra model parameters and negligible computational overhead (FSD can be as fast as greedy search). Despite the simplicity, FSD is surprisingly effective; Experiments show that FSD can outperform the canonical methods to date (i.e., nucleus sampling) as well as several strong baselines that were proposed recently. |
| title | A Frustratingly Simple Decoding Method for Neural Text Generation |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2305.12675 |