Prompting Science Report 2: The Decreasing Value of Chain of Thought in Prompting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meincke, Lennart, Mollick, Ethan, Mollick, Lilach, Shapiro, Dan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916784714547200
author Meincke, Lennart
Mollick, Ethan
Mollick, Lilach
Shapiro, Dan
author_facet Meincke, Lennart
Mollick, Ethan
Mollick, Lilach
Shapiro, Dan
contents This is the second in a series of short reports that seek to help business, education, and policy leaders understand the technical details of working with AI through rigorous testing. In this report, we investigate Chain-of-Thought (CoT) prompting, a technique that encourages a large language model (LLM) to "think step by step" (Wei et al., 2022). CoT is a widely adopted method for improving reasoning tasks, however, our findings reveal a more nuanced picture of its effectiveness. We demonstrate two things: - The effectiveness of Chain-of-Thought prompting can vary greatly depending on the type of task and model. For non-reasoning models, CoT generally improves average performance by a small amount, particularly if the model does not inherently engage in step-by-step processing by default. However, CoT can introduce more variability in answers, sometimes triggering occasional errors in questions the model would otherwise get right. We also found that many recent models perform some form of CoT reasoning even if not asked; for these models, a request to perform CoT had little impact. Performing CoT generally requires far more tokens (increasing cost and time) than direct answers. - For models designed with explicit reasoning capabilities, CoT prompting often results in only marginal, if any, gains in answer accuracy. However, it significantly increases the time and tokens needed to generate a response.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07142
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prompting Science Report 2: The Decreasing Value of Chain of Thought in Prompting
Meincke, Lennart
Mollick, Ethan
Mollick, Lilach
Shapiro, Dan
Computation and Language
Artificial Intelligence
This is the second in a series of short reports that seek to help business, education, and policy leaders understand the technical details of working with AI through rigorous testing. In this report, we investigate Chain-of-Thought (CoT) prompting, a technique that encourages a large language model (LLM) to "think step by step" (Wei et al., 2022). CoT is a widely adopted method for improving reasoning tasks, however, our findings reveal a more nuanced picture of its effectiveness. We demonstrate two things: - The effectiveness of Chain-of-Thought prompting can vary greatly depending on the type of task and model. For non-reasoning models, CoT generally improves average performance by a small amount, particularly if the model does not inherently engage in step-by-step processing by default. However, CoT can introduce more variability in answers, sometimes triggering occasional errors in questions the model would otherwise get right. We also found that many recent models perform some form of CoT reasoning even if not asked; for these models, a request to perform CoT had little impact. Performing CoT generally requires far more tokens (increasing cost and time) than direct answers. - For models designed with explicit reasoning capabilities, CoT prompting often results in only marginal, if any, gains in answer accuracy. However, it significantly increases the time and tokens needed to generate a response.
title Prompting Science Report 2: The Decreasing Value of Chain of Thought in Prompting
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.07142