(Why) Is My Prompt Getting Worse? Rethinking Regression Testing for Evolving LLM APIs
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Wanqin, Yang, Chenyang, Kästner, Christian |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts
by: Yang, Chenyang, et al.
Published: (2025)
by: Yang, Chenyang, et al.
Published: (2025)
What Is Wrong with My Model? Identifying Systematic Problems with Semantic Data Slicing
by: Yang, Chenyang, et al.
Published: (2024)
by: Yang, Chenyang, et al.
Published: (2024)
Octopus: On-device language model for function calling of software APIs
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
The Prompt Alchemist: Automated LLM-Tailored Prompt Optimization for Test Case Generation
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs
by: Lu, Yuxuan, et al.
Published: (2026)
by: Lu, Yuxuan, et al.
Published: (2026)
FASTRIC: Prompt Specification Language for Verifiable LLM Interactions
by: Jin, Wen-Long
Published: (2025)
by: Jin, Wen-Long
Published: (2025)
HarnessLLM: Automatic Testing Harness Generation via Reinforcement Learning
by: Liu, Yujian, et al.
Published: (2025)
by: Liu, Yujian, et al.
Published: (2025)
A Taxonomy of Prompt Defects in LLM Systems
by: Tian, Haoye, et al.
Published: (2025)
by: Tian, Haoye, et al.
Published: (2025)
Functional Consistency of LLM Code Embeddings: A Self-Evolving Data Synthesis Framework for Benchmarking
by: Li, Zhuohao, et al.
Published: (2025)
by: Li, Zhuohao, et al.
Published: (2025)
Show and Tell: Prompt Strategies for Style Control in Multi-Turn LLM Code Generation
by: Bohr, Jeremiah
Published: (2025)
by: Bohr, Jeremiah
Published: (2025)
Enabling Communication via APIs for Mainframe Applications
by: Kanvar, Vini, et al.
Published: (2024)
by: Kanvar, Vini, et al.
Published: (2024)
TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?
by: Ahmed, Toufique, et al.
Published: (2024)
by: Ahmed, Toufique, et al.
Published: (2024)
MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems
by: Guo, Guoxiang, et al.
Published: (2024)
by: Guo, Guoxiang, et al.
Published: (2024)
MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
How Toxic Can You Get? Search-based Toxicity Testing for Large Language Models
by: Corbo, Simone, et al.
Published: (2025)
by: Corbo, Simone, et al.
Published: (2025)
Interpretable Online Log Analysis Using Large Language Models with Prompt Strategies
by: Liu, Yilun, et al.
Published: (2023)
by: Liu, Yilun, et al.
Published: (2023)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
ProbeLLM: Automating Principled Diagnosis of LLM Failures
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
Dynamic Scaling of Unit Tests for Code Reward Modeling
by: Ma, Zeyao, et al.
Published: (2025)
by: Ma, Zeyao, et al.
Published: (2025)
LangGPT: Rethinking Structured Reusable Prompt Design Framework for LLMs from the Programming Language
by: Wang, Ming, et al.
Published: (2024)
by: Wang, Ming, et al.
Published: (2024)
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure
by: Yang, Zheyuan, et al.
Published: (2025)
by: Yang, Zheyuan, et al.
Published: (2025)
Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
UnitCoder: Scalable Iterative Code Synthesis with Unit Test Guidance
by: Ma, Yichuan, et al.
Published: (2025)
by: Ma, Yichuan, et al.
Published: (2025)
Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study
by: Li, Bowen, et al.
Published: (2024)
by: Li, Bowen, et al.
Published: (2024)
RethinkMCTS: Refining Erroneous Thoughts in Monte Carlo Tree Search for Code Generation
by: Li, Qingyao, et al.
Published: (2024)
by: Li, Qingyao, et al.
Published: (2024)
HateModerate: Testing Hate Speech Detectors against Content Moderation Policies
by: Zheng, Jiangrui, et al.
Published: (2023)
by: Zheng, Jiangrui, et al.
Published: (2023)
Evaluating and Achieving Controllable Code Completion in Code LLM
by: Zhang, Jiajun, et al.
Published: (2026)
by: Zhang, Jiajun, et al.
Published: (2026)
CodeContests+: High-Quality Test Case Generation for Competitive Programming
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Test Amplification for REST APIs via Single and Multi-Agent LLM Systems
by: Nooyens, Robbe, et al.
Published: (2025)
by: Nooyens, Robbe, et al.
Published: (2025)
Code Fingerprints: Disentangled Attribution of LLM-Generated Code
by: Guo, Jiaxun, et al.
Published: (2026)
by: Guo, Jiaxun, et al.
Published: (2026)
DiffuTester: Accelerating Unit Test Generation for Diffusion LLMs via Mining Structural Pattern
by: Yang, Lekang, et al.
Published: (2025)
by: Yang, Lekang, et al.
Published: (2025)
Evaluating LLM-Based Goal Extraction in Requirements Engineering: Prompting Strategies and Their Limitations
by: Arnaudo, Anna, et al.
Published: (2026)
by: Arnaudo, Anna, et al.
Published: (2026)
Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
by: Xia, Chunqiu Steven, et al.
Published: (2024)
by: Xia, Chunqiu Steven, et al.
Published: (2024)
PACE: Improving Prompt with Actor-Critic Editing for Large Language Model
by: Dong, Yihong, et al.
Published: (2023)
by: Dong, Yihong, et al.
Published: (2023)
TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
by: Liu, Steven, et al.
Published: (2026)
by: Liu, Steven, et al.
Published: (2026)
AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor
by: Yang, Shu, et al.
Published: (2026)
by: Yang, Shu, et al.
Published: (2026)
It Only Gets Worse: Revisiting DL-Based Vulnerability Detectors from a Practical Perspective
by: Wang, Yunqian, et al.
Published: (2025)
by: Wang, Yunqian, et al.
Published: (2025)
MEMCoder: Multi-dimensional Evolving Memory for Private-Library-Oriented Code Generation
by: Li, Mofei, et al.
Published: (2026)
by: Li, Mofei, et al.
Published: (2026)
Planning to Explore: Curiosity-Driven Planning for LLM Test Generation
by: Amayuelas, Alfonso, et al.
Published: (2026)
by: Amayuelas, Alfonso, et al.
Published: (2026)
Automating Computational Reproducibility in Social Science: Comparing Prompt-Based and Agent-Based Approaches
by: Shah, Syed Mehtab Hussain, et al.
Published: (2026)
by: Shah, Syed Mehtab Hussain, et al.
Published: (2026)
Similar Items
-
What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts
by: Yang, Chenyang, et al.
Published: (2025) -
What Is Wrong with My Model? Identifying Systematic Problems with Semantic Data Slicing
by: Yang, Chenyang, et al.
Published: (2024) -
Octopus: On-device language model for function calling of software APIs
by: Chen, Wei, et al.
Published: (2024) -
The Prompt Alchemist: Automated LLM-Tailored Prompt Optimization for Test Case Generation
by: Gao, Shuzheng, et al.
Published: (2025) -
Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs
by: Lu, Yuxuan, et al.
Published: (2026)