AI Eval Forge: Mixed-Check Regression Testing for LLM and Agent Workflows

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Author: Katta, Mukunda Rao
Format: Recurso digital
Language:English
Published: Zenodo 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866902008650268672
author Katta, Mukunda Rao
author_facet Katta, Mukunda Rao
contents Large-model and agent teams often need faster regression checks than broad benchmark suites can provide. This paper presents AI Eval Forge, a zero-dependency evaluation harness for mixed-check regression testing across LLM and agent workflows. The tool supports exact-match, substring, regex, token-F1, JSON validity, JSON field equality, citation coverage, and bounded custom-expression checks in a compact case format that works with JSON or JSONL. The contribution is not a new benchmark. It is a small, inspectable evaluation layer that helps teams compare runs, catch regressions, and summarize pass rate, score, cost, and latency without standing up a heavy evaluation stack. The paper describes the harness design, check model, reporting format, and practical role of mixed-check cases in real workflow testing. The artifact bundle is connected to the ai-eval-forge package and the public paper repository at https://github.com/MukundaKatta/ai-eval-forge-paper.
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_20044318
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle AI Eval Forge: Mixed-Check Regression Testing for LLM and Agent Workflows
Katta, Mukunda Rao
AI evaluation
LLM regression testing
agent evaluation
software testing
structured outputs
developer tooling
Large-model and agent teams often need faster regression checks than broad benchmark suites can provide. This paper presents AI Eval Forge, a zero-dependency evaluation harness for mixed-check regression testing across LLM and agent workflows. The tool supports exact-match, substring, regex, token-F1, JSON validity, JSON field equality, citation coverage, and bounded custom-expression checks in a compact case format that works with JSON or JSONL. The contribution is not a new benchmark. It is a small, inspectable evaluation layer that helps teams compare runs, catch regressions, and summarize pass rate, score, cost, and latency without standing up a heavy evaluation stack. The paper describes the harness design, check model, reporting format, and practical role of mixed-check cases in real workflow testing. The artifact bundle is connected to the ai-eval-forge package and the public paper repository at https://github.com/MukundaKatta/ai-eval-forge-paper.
title AI Eval Forge: Mixed-Check Regression Testing for LLM and Agent Workflows
topic AI evaluation
LLM regression testing
agent evaluation
software testing
structured outputs
developer tooling
url https://doi.org/10.5281/zenodo.20044318