Benchmarking LLM Agent Efficiency in Production Systems: An Observational Prospective Methodology

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Authors: Barcelos Costa, Cleber, Claude, Sonnet 4.6
Format: Recurso digital
Published: Zenodo 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866901473024016384
author Barcelos Costa, Cleber
Claude, Sonnet 4.6
author_facet Barcelos Costa, Cleber
Claude, Sonnet 4.6
contents <p>Existing large language model (LLM) benchmarks measure model capability on synthetic tasks, but none address the operational efficiency of multi-agent LLM systems executing real production work over sustained sessions. This paper introduces an observational prospective methodology for benchmarking production LLM agent efficiency and applies it to a complete, instrumented production session (2026-04-03, 4.1 hours, Gallora ecosystem). We report the first end-to-end token accounting of a multi-agent session: 64,853,375 effective tokens processed at a 94.2% cache hit rate, producing 19 artifacts at a cost of $36.74 USD ($1.93 per artifact). We propose a standardized suite of six operational metrics — Cache Hit Rate (CHR), Output Density (OD), Agent Cost Multiplier (ACM), Cost Per Artifact (CPA), Tool Execution Ratio (TER), and Turns Per Hour (TPH) — as a reproducible benchmark framework for production agentic systems. Key finding: in production multi-agent systems, cost is dominated by context complexity (93.7% cache reads), not task complexity — a result with significant architectural implications for system design and cost governance.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19432547
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Benchmarking LLM Agent Efficiency in Production Systems: An Observational Prospective Methodology
Barcelos Costa, Cleber
Claude, Sonnet 4.6
LLM benchmarking
multi-agent systems
production observability
token economics
agent efficiency
cost measurement
prompt caching
cache hit rate
Claude Code
agentic systems
operational metrics
cost per artifact
<p>Existing large language model (LLM) benchmarks measure model capability on synthetic tasks, but none address the operational efficiency of multi-agent LLM systems executing real production work over sustained sessions. This paper introduces an observational prospective methodology for benchmarking production LLM agent efficiency and applies it to a complete, instrumented production session (2026-04-03, 4.1 hours, Gallora ecosystem). We report the first end-to-end token accounting of a multi-agent session: 64,853,375 effective tokens processed at a 94.2% cache hit rate, producing 19 artifacts at a cost of $36.74 USD ($1.93 per artifact). We propose a standardized suite of six operational metrics — Cache Hit Rate (CHR), Output Density (OD), Agent Cost Multiplier (ACM), Cost Per Artifact (CPA), Tool Execution Ratio (TER), and Turns Per Hour (TPH) — as a reproducible benchmark framework for production agentic systems. Key finding: in production multi-agent systems, cost is dominated by context complexity (93.7% cache reads), not task complexity — a result with significant architectural implications for system design and cost governance.</p>
title Benchmarking LLM Agent Efficiency in Production Systems: An Observational Prospective Methodology
topic LLM benchmarking
multi-agent systems
production observability
token economics
agent efficiency
cost measurement
prompt caching
cache hit rate
Claude Code
agentic systems
operational metrics
cost per artifact
url https://doi.org/10.5281/zenodo.19432547