Structured Prompts Improve Evaluation of Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aali, Asad, Mohsin, Muhammad Ahmed, Bikia, Vasiliki, Singhvi, Arnav, Gaus, Richard, Bedi, Suhana, Cui, Hejie, Fuentes, Miguel, Unell, Alyssa, Mai, Yifan, Cahoon, Jordan, Pfeffer, Michael, Daneshjou, Roxana, Koyejo, Sanmi, Alsentzer, Emily, Potts, Christopher, Shah, Nigam H., Chaudhari, Akshay S.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917376557056000
author Aali, Asad
Mohsin, Muhammad Ahmed
Bikia, Vasiliki
Singhvi, Arnav
Gaus, Richard
Bedi, Suhana
Cui, Hejie
Fuentes, Miguel
Unell, Alyssa
Mai, Yifan
Cahoon, Jordan
Pfeffer, Michael
Daneshjou, Roxana
Koyejo, Sanmi
Alsentzer, Emily
Potts, Christopher
Shah, Nigam H.
Chaudhari, Akshay S.
author_facet Aali, Asad
Mohsin, Muhammad Ahmed
Bikia, Vasiliki
Singhvi, Arnav
Gaus, Richard
Bedi, Suhana
Cui, Hejie
Fuentes, Miguel
Unell, Alyssa
Mai, Yifan
Cahoon, Jordan
Pfeffer, Michael
Daneshjou, Roxana
Koyejo, Sanmi
Alsentzer, Emily
Potts, Christopher
Shah, Nigam H.
Chaudhari, Akshay S.
contents As language models (LMs) are increasingly adopted across domains, high-quality benchmarking frameworks are essential for guiding deployment decisions. In practice, however, frameworks such as Holistic Evaluation of Language Models (HELM) typically evaluate models under a single static prompt configuration, even though model behavior depends strongly on prompt choice. As a result, reported scores can reflect prompt choice as much as model capability. Declarative prompting frameworks such as DSPy offer a scalable way to evaluate models under a set of structured prompting strategies rather than a static prompt configuration. We present a reproducible DSPy+HELM framework for studying how prompt choice impacts reported benchmark outcomes. Using five prompting methods, we evaluate four frontier and two open-source LMs across seven benchmarks against existing HELM baseline scores. By evaluating LMs across a family of prompt configurations, we find that prompt choice can materially impact leaderboard outcomes. In particular, structured prompting improves performance (by 6% on average), alters comparisons (leaderboard rankings shift on 5/7 benchmarks), with most gains coming from introducing chain-of-thought, and little additional benefit from more advanced optimizers. To our knowledge, this is the first study to systematically integrate structured prompting into an established evaluation framework and quantify how prompt choice alone can impact benchmark conclusions. We open-source (i) DSPy+HELM Evaluation (https://github.com/stanford-crfm/helm/pull/3893) and (ii) Prompt Optimization Pipeline (https://github.com/StanfordMIMI/dspy-helm).
format Preprint
id arxiv_https___arxiv_org_abs_2511_20836
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Structured Prompts Improve Evaluation of Language Models
Aali, Asad
Mohsin, Muhammad Ahmed
Bikia, Vasiliki
Singhvi, Arnav
Gaus, Richard
Bedi, Suhana
Cui, Hejie
Fuentes, Miguel
Unell, Alyssa
Mai, Yifan
Cahoon, Jordan
Pfeffer, Michael
Daneshjou, Roxana
Koyejo, Sanmi
Alsentzer, Emily
Potts, Christopher
Shah, Nigam H.
Chaudhari, Akshay S.
Computation and Language
Artificial Intelligence
Machine Learning
As language models (LMs) are increasingly adopted across domains, high-quality benchmarking frameworks are essential for guiding deployment decisions. In practice, however, frameworks such as Holistic Evaluation of Language Models (HELM) typically evaluate models under a single static prompt configuration, even though model behavior depends strongly on prompt choice. As a result, reported scores can reflect prompt choice as much as model capability. Declarative prompting frameworks such as DSPy offer a scalable way to evaluate models under a set of structured prompting strategies rather than a static prompt configuration. We present a reproducible DSPy+HELM framework for studying how prompt choice impacts reported benchmark outcomes. Using five prompting methods, we evaluate four frontier and two open-source LMs across seven benchmarks against existing HELM baseline scores. By evaluating LMs across a family of prompt configurations, we find that prompt choice can materially impact leaderboard outcomes. In particular, structured prompting improves performance (by 6% on average), alters comparisons (leaderboard rankings shift on 5/7 benchmarks), with most gains coming from introducing chain-of-thought, and little additional benefit from more advanced optimizers. To our knowledge, this is the first study to systematically integrate structured prompting into an established evaluation framework and quantify how prompt choice alone can impact benchmark conclusions. We open-source (i) DSPy+HELM Evaluation (https://github.com/stanford-crfm/helm/pull/3893) and (ii) Prompt Optimization Pipeline (https://github.com/StanfordMIMI/dspy-helm).
title Structured Prompts Improve Evaluation of Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2511.20836