What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hedderich, Michael A., Wang, Anyi, Zhao, Raoyuan, Eichin, Florian, Fischer, Jonas, Plank, Barbara
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909628880650240
author Hedderich, Michael A.
Wang, Anyi
Zhao, Raoyuan
Eichin, Florian
Fischer, Jonas
Plank, Barbara
author_facet Hedderich, Michael A.
Wang, Anyi
Zhao, Raoyuan
Eichin, Florian
Fischer, Jonas
Plank, Barbara
contents Prompt engineering for large language models is challenging, as even small prompt perturbations or model changes can significantly impact the generated output texts. Existing evaluation methods of LLM outputs, either automated metrics or human evaluation, have limitations, such as providing limited insights or being labor-intensive. We propose Spotlight, a new approach that combines both automation and human analysis. Based on data mining techniques, we automatically distinguish between random (decoding) variations and systematic differences in language model outputs. This process provides token patterns that describe the systematic differences and guide the user in manually analyzing the effects of their prompts and changes in models efficiently. We create three benchmarks to quantitatively test the reliability of token pattern extraction methods and demonstrate that our approach provides new insights into established prompt data. From a human-centric perspective, through demonstration studies and a user study, we show that our token pattern approach helps users understand the systematic differences of language model outputs. We are further able to discover relevant differences caused by prompt and model changes (e.g. related to gender or culture), thus supporting the prompt engineering process and human-centric model behavior research.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15815
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns
Hedderich, Michael A.
Wang, Anyi
Zhao, Raoyuan
Eichin, Florian
Fischer, Jonas
Plank, Barbara
Computation and Language
Human-Computer Interaction
Machine Learning
Prompt engineering for large language models is challenging, as even small prompt perturbations or model changes can significantly impact the generated output texts. Existing evaluation methods of LLM outputs, either automated metrics or human evaluation, have limitations, such as providing limited insights or being labor-intensive. We propose Spotlight, a new approach that combines both automation and human analysis. Based on data mining techniques, we automatically distinguish between random (decoding) variations and systematic differences in language model outputs. This process provides token patterns that describe the systematic differences and guide the user in manually analyzing the effects of their prompts and changes in models efficiently. We create three benchmarks to quantitatively test the reliability of token pattern extraction methods and demonstrate that our approach provides new insights into established prompt data. From a human-centric perspective, through demonstration studies and a user study, we show that our token pattern approach helps users understand the systematic differences of language model outputs. We are further able to discover relevant differences caused by prompt and model changes (e.g. related to gender or culture), thus supporting the prompt engineering process and human-centric model behavior research.
title What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns
topic Computation and Language
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2504.15815