To Burst or Not to Burst: Generating and Quantifying Improbable Text

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sasse, Kuleen, Barham, Samuel, Kayi, Efsun Sarioglu, Staley, Edward W.
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914655886114816
author Sasse, Kuleen
Barham, Samuel
Kayi, Efsun Sarioglu
Staley, Edward W.
author_facet Sasse, Kuleen
Barham, Samuel
Kayi, Efsun Sarioglu
Staley, Edward W.
contents While large language models (LLMs) are extremely capable at text generation, their outputs are still distinguishable from human-authored text. We explore this separation across many metrics over text, many sampling techniques, many types of text data, and across two popular LLMs, LLaMA and Vicuna. Along the way, we introduce a new metric, recoverability, to highlight differences between human and machine text; and we propose a new sampling technique, burst sampling, designed to close this gap. We find that LLaMA and Vicuna have distinct distributions under many of the metrics, and that this influences our results: Recoverability separates real from fake text better than any other metric when using LLaMA. When using Vicuna, burst sampling produces text which is distributionally closer to real text compared to other sampling techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2401_15476
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle To Burst or Not to Burst: Generating and Quantifying Improbable Text
Sasse, Kuleen
Barham, Samuel
Kayi, Efsun Sarioglu
Staley, Edward W.
Computation and Language
While large language models (LLMs) are extremely capable at text generation, their outputs are still distinguishable from human-authored text. We explore this separation across many metrics over text, many sampling techniques, many types of text data, and across two popular LLMs, LLaMA and Vicuna. Along the way, we introduce a new metric, recoverability, to highlight differences between human and machine text; and we propose a new sampling technique, burst sampling, designed to close this gap. We find that LLaMA and Vicuna have distinct distributions under many of the metrics, and that this influences our results: Recoverability separates real from fake text better than any other metric when using LLaMA. When using Vicuna, burst sampling produces text which is distributionally closer to real text compared to other sampling techniques.
title To Burst or Not to Burst: Generating and Quantifying Improbable Text
topic Computation and Language
url https://arxiv.org/abs/2401.15476