Continuous Benchmark Generation for Evaluating Enterprise-scale LLM Agents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Saxena, Divyanshu, Maurya, Rishikesh, Ou, Xiaoxuan, Somashekar, Gagan, Gupta, Shachee Mishra, Iyer, Arun, Kang, Yu, Bansal, Chetan, Akella, Aditya, Rajmohan, Saravan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918200420073472
author Saxena, Divyanshu
Maurya, Rishikesh
Ou, Xiaoxuan
Somashekar, Gagan
Gupta, Shachee Mishra
Iyer, Arun
Kang, Yu
Bansal, Chetan
Akella, Aditya
Rajmohan, Saravan
author_facet Saxena, Divyanshu
Maurya, Rishikesh
Ou, Xiaoxuan
Somashekar, Gagan
Gupta, Shachee Mishra
Iyer, Arun
Kang, Yu
Bansal, Chetan
Akella, Aditya
Rajmohan, Saravan
contents The rapid adoption of AI agents across domains has made systematic evaluation crucial for ensuring their usefulness and successful production deployment. Evaluation of AI agents typically involves using a fixed set of benchmarks and computing multiple evaluation metrics for the agent. While sufficient for simple coding tasks, these benchmarks fall short for enterprise-scale agents, where services and requirements evolve continuously and ground-truth examples are sparse. We propose a process of benchmark generation that helps evolve the benchmarks as the requirements change and perform robust evaluation of evolving AI agents. We instantiate this approach for a case study of service migration from one deployment platform to another at a large public enterprise. Our approach relies on semi-structured documents where developers express the high-level intent, and uses state-of-the-art LLMs to generate benchmarks from just a small number of such documents. Overall, this process results in a maintainable evaluation framework, enabling rapid feedback on agent performance and facilitating targeted improvements.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10049
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Continuous Benchmark Generation for Evaluating Enterprise-scale LLM Agents
Saxena, Divyanshu
Maurya, Rishikesh
Ou, Xiaoxuan
Somashekar, Gagan
Gupta, Shachee Mishra
Iyer, Arun
Kang, Yu
Bansal, Chetan
Akella, Aditya
Rajmohan, Saravan
Software Engineering
The rapid adoption of AI agents across domains has made systematic evaluation crucial for ensuring their usefulness and successful production deployment. Evaluation of AI agents typically involves using a fixed set of benchmarks and computing multiple evaluation metrics for the agent. While sufficient for simple coding tasks, these benchmarks fall short for enterprise-scale agents, where services and requirements evolve continuously and ground-truth examples are sparse. We propose a process of benchmark generation that helps evolve the benchmarks as the requirements change and perform robust evaluation of evolving AI agents. We instantiate this approach for a case study of service migration from one deployment platform to another at a large public enterprise. Our approach relies on semi-structured documents where developers express the high-level intent, and uses state-of-the-art LLMs to generate benchmarks from just a small number of such documents. Overall, this process results in a maintainable evaluation framework, enabling rapid feedback on agent performance and facilitating targeted improvements.
title Continuous Benchmark Generation for Evaluating Enterprise-scale LLM Agents
topic Software Engineering
url https://arxiv.org/abs/2511.10049