| _version_ | 1866901473024016384 |
|---|---|
| author | Barcelos Costa, Cleber Claude, Sonnet 4.6 |
| author_facet | Barcelos Costa, Cleber Claude, Sonnet 4.6 |
| contents | <p>Existing large language model (LLM) benchmarks measure model capability on synthetic tasks, but none address the operational efficiency of multi-agent LLM systems executing real production work over sustained sessions. This paper introduces an observational prospective methodology for benchmarking production LLM agent efficiency and applies it to a complete, instrumented production session (2026-04-03, 4.1 hours, Gallora ecosystem). We report the first end-to-end token accounting of a multi-agent session: 64,853,375 effective tokens processed at a 94.2% cache hit rate, producing 19 artifacts at a cost of $36.74 USD ($1.93 per artifact). We propose a standardized suite of six operational metrics — Cache Hit Rate (CHR), Output Density (OD), Agent Cost Multiplier (ACM), Cost Per Artifact (CPA), Tool Execution Ratio (TER), and Turns Per Hour (TPH) — as a reproducible benchmark framework for production agentic systems. Key finding: in production multi-agent systems, cost is dominated by context complexity (93.7% cache reads), not task complexity — a result with significant architectural implications for system design and cost governance.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19432547 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Benchmarking LLM Agent Efficiency in Production Systems: An Observational Prospective Methodology Barcelos Costa, Cleber Claude, Sonnet 4.6 LLM benchmarking multi-agent systems production observability token economics agent efficiency cost measurement prompt caching cache hit rate Claude Code agentic systems operational metrics cost per artifact <p>Existing large language model (LLM) benchmarks measure model capability on synthetic tasks, but none address the operational efficiency of multi-agent LLM systems executing real production work over sustained sessions. This paper introduces an observational prospective methodology for benchmarking production LLM agent efficiency and applies it to a complete, instrumented production session (2026-04-03, 4.1 hours, Gallora ecosystem). We report the first end-to-end token accounting of a multi-agent session: 64,853,375 effective tokens processed at a 94.2% cache hit rate, producing 19 artifacts at a cost of $36.74 USD ($1.93 per artifact). We propose a standardized suite of six operational metrics — Cache Hit Rate (CHR), Output Density (OD), Agent Cost Multiplier (ACM), Cost Per Artifact (CPA), Tool Execution Ratio (TER), and Turns Per Hour (TPH) — as a reproducible benchmark framework for production agentic systems. Key finding: in production multi-agent systems, cost is dominated by context complexity (93.7% cache reads), not task complexity — a result with significant architectural implications for system design and cost governance.</p> |
| title | Benchmarking LLM Agent Efficiency in Production Systems: An Observational Prospective Methodology |
| topic | LLM benchmarking multi-agent systems production observability token economics agent efficiency cost measurement prompt caching cache hit rate Claude Code agentic systems operational metrics cost per artifact |
| url | https://doi.org/10.5281/zenodo.19432547 |