Prompting Science Report 1: Prompt Engineering is Complicated and Contingent

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Meincke, Lennart, Mollick, Ethan, Mollick, Lilach, Shapiro, Dan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910862244052992
author Meincke, Lennart
Mollick, Ethan
Mollick, Lilach
Shapiro, Dan
author_facet Meincke, Lennart
Mollick, Ethan
Mollick, Lilach
Shapiro, Dan
contents This is the first of a series of short reports that seek to help business, education, and policy leaders understand the technical details of working with AI through rigorous testing. In this report, we demonstrate two things: - There is no single standard for measuring whether a Large Language Model (LLM) passes a benchmark, and that choosing a standard has a big impact on how well the LLM does on that benchmark. The standard you choose will depend on your goals for using an LLM in a particular case. - It is hard to know in advance whether a particular prompting approach will help or harm the LLM's ability to answer any particular question. Specifically, we find that sometimes being polite to the LLM helps performance, and sometimes it lowers performance. We also find that constraining the AI's answers helps performance in some cases, though it may lower performance in other cases. Taken together, this suggests that benchmarking AI performance is not one-size-fits-all, and also that particular prompting formulas or approaches, like being polite to the AI, are not universally valuable.
format Preprint
id arxiv_https___arxiv_org_abs_2503_04818
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prompting Science Report 1: Prompt Engineering is Complicated and Contingent
Meincke, Lennart
Mollick, Ethan
Mollick, Lilach
Shapiro, Dan
Computation and Language
Artificial Intelligence
This is the first of a series of short reports that seek to help business, education, and policy leaders understand the technical details of working with AI through rigorous testing. In this report, we demonstrate two things: - There is no single standard for measuring whether a Large Language Model (LLM) passes a benchmark, and that choosing a standard has a big impact on how well the LLM does on that benchmark. The standard you choose will depend on your goals for using an LLM in a particular case. - It is hard to know in advance whether a particular prompting approach will help or harm the LLM's ability to answer any particular question. Specifically, we find that sometimes being polite to the LLM helps performance, and sometimes it lowers performance. We also find that constraining the AI's answers helps performance in some cases, though it may lower performance in other cases. Taken together, this suggests that benchmarking AI performance is not one-size-fits-all, and also that particular prompting formulas or approaches, like being polite to the AI, are not universally valuable.
title Prompting Science Report 1: Prompt Engineering is Complicated and Contingent
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.04818