Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Wei, Yang, Yixiao, Hu, Qiang, Ying, Shi, Jin, Zhi, Du, Bo, Xing, Zhenchang, Li, Tianlin, Shi, Junjie, Liu, Yang, Jiang, Linxiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908508286353408
author Ma, Wei
Yang, Yixiao
Hu, Qiang
Ying, Shi
Jin, Zhi
Du, Bo
Xing, Zhenchang
Li, Tianlin
Shi, Junjie
Liu, Yang
Jiang, Linxiao
author_facet Ma, Wei
Yang, Yixiao
Hu, Qiang
Ying, Shi
Jin, Zhi
Du, Bo
Xing, Zhenchang
Li, Tianlin
Shi, Junjie
Liu, Yang
Jiang, Linxiao
contents Applications of Large Language Models~(LLMs) have evolved from simple text generators into complex software systems that integrate retrieval augmentation, tool invocation, and multi-turn interactions. Their inherent non-determinism, dynamism, and context dependence pose fundamental challenges for quality assurance. This paper decomposes LLM applications into a three-layer architecture: \textbf{\textit{System Shell Layer}}, \textbf{\textit{Prompt Orchestration Layer}}, and \textbf{\textit{LLM Inference Core}}. We then assess the applicability of traditional software testing methods in each layer: directly applicable at the shell layer, requiring semantic reinterpretation at the orchestration layer, and necessitating paradigm shifts at the inference core. A comparative analysis of Testing AI methods from the software engineering community and safety analysis techniques from the AI community reveals structural disconnects in testing unit abstraction, evaluation metrics, and lifecycle management. We identify four fundamental differences that underlie 6 core challenges. To address these, we propose four types of collaborative strategies (\emph{Retain}, \emph{Translate}, \emph{Integrate}, and \emph{Runtime}) and explore a closed-loop, trustworthy quality assurance framework that combines pre-deployment validation with runtime monitoring. Based on these strategies, we offer practical guidance and a protocol proposal to support the standardization and tooling of LLM application testing. We propose a protocol \textbf{\textit{Agent Interaction Communication Language}} (AICL) that is used to communicate between AI agents. AICL has the test-oriented features and is easily integrated in the current agent framework.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20737
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol
Ma, Wei
Yang, Yixiao
Hu, Qiang
Ying, Shi
Jin, Zhi
Du, Bo
Xing, Zhenchang
Li, Tianlin
Shi, Junjie
Liu, Yang
Jiang, Linxiao
Software Engineering
Artificial Intelligence
Applications of Large Language Models~(LLMs) have evolved from simple text generators into complex software systems that integrate retrieval augmentation, tool invocation, and multi-turn interactions. Their inherent non-determinism, dynamism, and context dependence pose fundamental challenges for quality assurance. This paper decomposes LLM applications into a three-layer architecture: \textbf{\textit{System Shell Layer}}, \textbf{\textit{Prompt Orchestration Layer}}, and \textbf{\textit{LLM Inference Core}}. We then assess the applicability of traditional software testing methods in each layer: directly applicable at the shell layer, requiring semantic reinterpretation at the orchestration layer, and necessitating paradigm shifts at the inference core. A comparative analysis of Testing AI methods from the software engineering community and safety analysis techniques from the AI community reveals structural disconnects in testing unit abstraction, evaluation metrics, and lifecycle management. We identify four fundamental differences that underlie 6 core challenges. To address these, we propose four types of collaborative strategies (\emph{Retain}, \emph{Translate}, \emph{Integrate}, and \emph{Runtime}) and explore a closed-loop, trustworthy quality assurance framework that combines pre-deployment validation with runtime monitoring. Based on these strategies, we offer practical guidance and a protocol proposal to support the standardization and tooling of LLM application testing. We propose a protocol \textbf{\textit{Agent Interaction Communication Language}} (AICL) that is used to communicate between AI agents. AICL has the test-oriented features and is easily integrated in the current agent framework.
title Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2508.20737