Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Wei, Yang, Yixiao, Hu, Qiang, Ying, Shi, Jin, Zhi, Du, Bo, Xing, Zhenchang, Li, Tianlin, Shi, Junjie, Liu, Yang, Jiang, Linxiao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compiling Code LLMs into Lightweight Executables
by: Shi, Jieke, et al.
Published: (2026)
by: Shi, Jieke, et al.
Published: (2026)
Editorial for the special issue on software refactoring: Application breadth and technical depth
by: Zhenchang Xing
Published: (2024)
by: Zhenchang Xing
Published: (2024)
Knowledge-Guided Multi-Agent Framework for Application-Level Software Code Generation
by: Xiong, Qian, et al.
Published: (2025)
by: Xiong, Qian, et al.
Published: (2025)
Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents
by: Chen, Zhi, et al.
Published: (2026)
by: Chen, Zhi, et al.
Published: (2026)
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
Automated Soap Opera Testing Directed by LLMs and Scenario Knowledge: Feasibility, Challenges, and Road Ahead
by: Su, Yanqi, et al.
Published: (2024)
by: Su, Yanqi, et al.
Published: (2024)
LogiAgent: Automated Logical Testing for REST Systems with LLM-Based Multi-Agents
by: Zhang, Ke, et al.
Published: (2025)
by: Zhang, Ke, et al.
Published: (2025)
A Comparative Study of Android Performance Issues in Real-world Applications and Literature
by: Liao, Dianshu, et al.
Published: (2024)
by: Liao, Dianshu, et al.
Published: (2024)
Exploring the Power of Diffusion Large Language Models for Software Engineering: An Empirical Investigation
by: Zhang, Jingyao, et al.
Published: (2025)
by: Zhang, Jingyao, et al.
Published: (2025)
PTMPicker: Facilitating Efficient Pretrained Model Selection for Application Developers
by: Liu, Pei, et al.
Published: (2025)
by: Liu, Pei, et al.
Published: (2025)
Rethinking Technology Stack Selection with AI Coding Proficiency
by: Zhang, Xiaoyu, et al.
Published: (2025)
by: Zhang, Xiaoyu, et al.
Published: (2025)
A Solution toward Transparent and Practical AI Regulation: Privacy Nutrition Labels for Open-source Generative AI-based Applications
by: Si, Meixue, et al.
Published: (2024)
by: Si, Meixue, et al.
Published: (2024)
From Code to Courtroom: LLMs as the New Software Judges
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
REprompt: Prompt Generation for Intelligent Software Development Guided by Requirements Engineering
by: Shi, Junjie, et al.
Published: (2026)
by: Shi, Junjie, et al.
Published: (2026)
MeTMaP: Metamorphic Testing for Detecting False Vector Matching Problems in LLM Augmented Generation
by: Wang, Guanyu, et al.
Published: (2024)
by: Wang, Guanyu, et al.
Published: (2024)
Human-Like Code Quality Evaluation through LLM-based Recursive Semantic Comprehension
by: Xu, Fangzhou, et al.
Published: (2024)
by: Xu, Fangzhou, et al.
Published: (2024)
When Prompt Engineering Meets Software Engineering: CNL-P as Natural and Robust "APIs'' for Human-AI Interaction
by: Xing, Zhenchang, et al.
Published: (2025)
by: Xing, Zhenchang, et al.
Published: (2025)
(Why) Is My Prompt Getting Worse? Rethinking Regression Testing for Evolving LLM APIs
by: Ma, Wanqin, et al.
Published: (2023)
by: Ma, Wanqin, et al.
Published: (2023)
Explore-Construct-Filter: An Automated Framework for Rich and Reliable API Knowledge Graph Construction
by: Sun, Yanbang, et al.
Published: (2025)
by: Sun, Yanbang, et al.
Published: (2025)
iPanda: An LLM-based Agent for Automated Conformance Testing of Communication Protocols
by: Sun, Xikai, et al.
Published: (2025)
by: Sun, Xikai, et al.
Published: (2025)
Uncovering Business Logic Bugs via Semantics-Driven Unit Test Generation
by: Yang, Chen, et al.
Published: (2026)
by: Yang, Chen, et al.
Published: (2026)
Hallucination Detection for LLM-based Text-to-SQL Generation via Two-Stage Metamorphic Testing
by: Yang, Bo, et al.
Published: (2025)
by: Yang, Bo, et al.
Published: (2025)
Navigating the Labyrinth: Path-Sensitive Unit Test Generation with Large Language Models
by: Liao, Dianshu, et al.
Published: (2025)
by: Liao, Dianshu, et al.
Published: (2025)
Large Language Models for Unit Test Generation: Achievements, Challenges, and Opportunities
by: Chu, Bei, et al.
Published: (2025)
by: Chu, Bei, et al.
Published: (2025)
Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak Attacks
by: Zhou, Shide, et al.
Published: (2024)
by: Zhou, Shide, et al.
Published: (2024)
CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges
by: Zhang, Kechi, et al.
Published: (2024)
by: Zhang, Kechi, et al.
Published: (2024)
Evaluation-Driven Development and Operations of LLM Agents: A Process Model and Reference Architecture
by: Xia, Boming, et al.
Published: (2024)
by: Xia, Boming, et al.
Published: (2024)
A Microservice Graph Generator with Production Characteristics
by: Du, Fanrong, et al.
Published: (2024)
by: Du, Fanrong, et al.
Published: (2024)
Back to the Basics: Rethinking Issue-Commit Linking with LLM-Assisted Retrieval
by: Huang, Huihui, et al.
Published: (2025)
by: Huang, Huihui, et al.
Published: (2025)
HITS: High-coverage LLM-based Unit Test Generation via Method Slicing
by: Wang, Zejun, et al.
Published: (2024)
by: Wang, Zejun, et al.
Published: (2024)
CodeChemist: Functional Knowledge Transfer for Low-Resource Code Generation via Test-Time Scaling
by: Wang, Kaixin, et al.
Published: (2025)
by: Wang, Kaixin, et al.
Published: (2025)
The Current Challenges of Software Engineering in the Era of Large Language Models
by: Gao, Cuiyun, et al.
Published: (2024)
by: Gao, Cuiyun, et al.
Published: (2024)
Mut4All: Fuzzing Compilers via LLM-Synthesized Mutators Learned from Bug Reports
by: Wang, Bo, et al.
Published: (2025)
by: Wang, Bo, et al.
Published: (2025)
Automated Static Warning Identification via Path-based Semantic Representation
by: Zhang, Yuwei, et al.
Published: (2023)
by: Zhang, Yuwei, et al.
Published: (2023)
STEAM: Simulating the InTeractive BEhavior of ProgrAMmers for Automatic Bug Fixing
by: Zhang, Yuwei, et al.
Published: (2023)
by: Zhang, Yuwei, et al.
Published: (2023)
Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
Reflective Unit Test Generation for Precise Type Error Detection with Large Language Models
by: Yang, Chen, et al.
Published: (2025)
by: Yang, Chen, et al.
Published: (2025)
QuanTest: Entanglement-Guided Testing of Quantum Neural Network Systems
by: Shi, Jinjing, et al.
Published: (2024)
by: Shi, Jinjing, et al.
Published: (2024)
Test vs Mutant: Adversarial LLM Agents for Robust Unit Test Generation
by: Chang, Pengyu, et al.
Published: (2026)
by: Chang, Pengyu, et al.
Published: (2026)
When AI Takes the Wheel: Security Analysis of Framework-Constrained Program Generation
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
Similar Items
-
Compiling Code LLMs into Lightweight Executables
by: Shi, Jieke, et al.
Published: (2026) -
Editorial for the special issue on software refactoring: Application breadth and technical depth
by: Zhenchang Xing
Published: (2024) -
Knowledge-Guided Multi-Agent Framework for Application-Level Software Code Generation
by: Xiong, Qian, et al.
Published: (2025) -
Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents
by: Chen, Zhi, et al.
Published: (2026) -
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
by: He, Junda, et al.
Published: (2025)