(Why) Is My Prompt Getting Worse? Rethinking Regression Testing for Evolving LLM APIs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Wanqin, Yang, Chenyang, Kästner, Christian
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917583809150976
author Ma, Wanqin
Yang, Chenyang
Kästner, Christian
author_facet Ma, Wanqin
Yang, Chenyang
Kästner, Christian
contents Large Language Models (LLMs) are increasingly integrated into software applications. Downstream application developers often access LLMs through APIs provided as a service. However, LLM APIs are often updated silently and scheduled to be deprecated, forcing users to continuously adapt to evolving models. This can cause performance regression and affect prompt design choices, as evidenced by our case study on toxicity detection. Based on our case study, we emphasize the need for and re-examine the concept of regression testing for evolving LLM APIs. We argue that regression testing LLMs requires fundamental changes to traditional testing approaches, due to different correctness notions, prompting brittleness, and non-determinism in LLM APIs.
format Preprint
id arxiv_https___arxiv_org_abs_2311_11123
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle (Why) Is My Prompt Getting Worse? Rethinking Regression Testing for Evolving LLM APIs
Ma, Wanqin
Yang, Chenyang
Kästner, Christian
Software Engineering
Computation and Language
Large Language Models (LLMs) are increasingly integrated into software applications. Downstream application developers often access LLMs through APIs provided as a service. However, LLM APIs are often updated silently and scheduled to be deprecated, forcing users to continuously adapt to evolving models. This can cause performance regression and affect prompt design choices, as evidenced by our case study on toxicity detection. Based on our case study, we emphasize the need for and re-examine the concept of regression testing for evolving LLM APIs. We argue that regression testing LLMs requires fundamental changes to traditional testing approaches, due to different correctness notions, prompting brittleness, and non-determinism in LLM APIs.
title (Why) Is My Prompt Getting Worse? Rethinking Regression Testing for Evolving LLM APIs
topic Software Engineering
Computation and Language
url https://arxiv.org/abs/2311.11123