Strings from the Library of Babel: Random Sampling as a Strong Baseline for Prompt Optimisation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Yao, Wang, Jiayi, Tang, Raphael, Riedel, Sebastian, Stenetorp, Pontus
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909172222656512
author Lu, Yao
Wang, Jiayi
Tang, Raphael
Riedel, Sebastian
Stenetorp, Pontus
author_facet Lu, Yao
Wang, Jiayi
Tang, Raphael
Riedel, Sebastian
Stenetorp, Pontus
contents Recent prompt optimisation approaches use the generative nature of language models to produce prompts -- even rivaling the performance of human-curated prompts. In this paper, we demonstrate that randomly sampling tokens from the model vocabulary as ``separators'' can be as effective as language models for prompt-style text classification. Our experiments show that random separators are competitive baselines, having less than a 1% difference compared to previous self-optimisation methods and showing a 12% average relative improvement over strong human baselines across nine text classification tasks and eight language models. We further analyse this phenomenon in detail using three different random generation strategies, establishing that the language space is rich with potentially good separators, with a greater than 40% average chance that a randomly drawn separator performs better than human-curated separators. These observations challenge the common assumption that an effective prompt should be human readable or task relevant and establish a strong baseline for prompt optimisation research.
format Preprint
id arxiv_https___arxiv_org_abs_2311_09569
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Strings from the Library of Babel: Random Sampling as a Strong Baseline for Prompt Optimisation
Lu, Yao
Wang, Jiayi
Tang, Raphael
Riedel, Sebastian
Stenetorp, Pontus
Computation and Language
Artificial Intelligence
Recent prompt optimisation approaches use the generative nature of language models to produce prompts -- even rivaling the performance of human-curated prompts. In this paper, we demonstrate that randomly sampling tokens from the model vocabulary as ``separators'' can be as effective as language models for prompt-style text classification. Our experiments show that random separators are competitive baselines, having less than a 1% difference compared to previous self-optimisation methods and showing a 12% average relative improvement over strong human baselines across nine text classification tasks and eight language models. We further analyse this phenomenon in detail using three different random generation strategies, establishing that the language space is rich with potentially good separators, with a greater than 40% average chance that a randomly drawn separator performs better than human-curated separators. These observations challenge the common assumption that an effective prompt should be human readable or task relevant and establish a strong baseline for prompt optimisation research.
title Strings from the Library of Babel: Random Sampling as a Strong Baseline for Prompt Optimisation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2311.09569