Prompting Science Report 3: I'll pay you or I'll kill you -- but will you care?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meincke, Lennart, Mollick, Ethan, Mollick, Lilach, Shapiro, Dan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916875239161856
author Meincke, Lennart
Mollick, Ethan
Mollick, Lilach
Shapiro, Dan
author_facet Meincke, Lennart
Mollick, Ethan
Mollick, Lilach
Shapiro, Dan
contents This is the third in a series of short reports that seek to help business, education, and policy leaders understand the technical details of working with AI through rigorous testing. In this report, we investigate two commonly held prompting beliefs: a) offering to tip the AI model and b) threatening the AI model. Tipping was a commonly shared tactic for improving AI performance and threats have been endorsed by Google Founder Sergey Brin (All-In, May 2025, 8:20) who observed that 'models tend to do better if you threaten them,' a claim we subject to empirical testing here. We evaluate model performance on GPQA (Rein et al. 2024) and MMLU-Pro (Wang et al. 2024). We demonstrate two things: - Threatening or tipping a model generally has no significant effect on benchmark performance. - Prompt variations can significantly affect performance on a per-question level. However, it is hard to know in advance whether a particular prompting approach will help or harm the LLM's ability to answer any particular question. Taken together, this suggests that simple prompting variations might not be as effective as previously assumed, especially for difficult problems. However, as reported previously (Meincke et al. 2025a), prompting approaches can yield significantly different results for individual questions.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00614
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prompting Science Report 3: I'll pay you or I'll kill you -- but will you care?
Meincke, Lennart
Mollick, Ethan
Mollick, Lilach
Shapiro, Dan
Computation and Language
Artificial Intelligence
This is the third in a series of short reports that seek to help business, education, and policy leaders understand the technical details of working with AI through rigorous testing. In this report, we investigate two commonly held prompting beliefs: a) offering to tip the AI model and b) threatening the AI model. Tipping was a commonly shared tactic for improving AI performance and threats have been endorsed by Google Founder Sergey Brin (All-In, May 2025, 8:20) who observed that 'models tend to do better if you threaten them,' a claim we subject to empirical testing here. We evaluate model performance on GPQA (Rein et al. 2024) and MMLU-Pro (Wang et al. 2024). We demonstrate two things: - Threatening or tipping a model generally has no significant effect on benchmark performance. - Prompt variations can significantly affect performance on a per-question level. However, it is hard to know in advance whether a particular prompting approach will help or harm the LLM's ability to answer any particular question. Taken together, this suggests that simple prompting variations might not be as effective as previously assumed, especially for difficult problems. However, as reported previously (Meincke et al. 2025a), prompting approaches can yield significantly different results for individual questions.
title Prompting Science Report 3: I'll pay you or I'll kill you -- but will you care?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.00614