E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sadhuka, Shuvom, Prinster, Drew, Fannjiang, Clara, Scalia, Gabriele, Berger, Bonnie, Regev, Aviv, Wang, Hanchen
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914613559296000
author Sadhuka, Shuvom
Prinster, Drew
Fannjiang, Clara
Scalia, Gabriele
Berger, Bonnie
Regev, Aviv
Wang, Hanchen
author_facet Sadhuka, Shuvom
Prinster, Drew
Fannjiang, Clara
Scalia, Gabriele
Berger, Bonnie
Regev, Aviv
Wang, Hanchen
contents Agentic AI systems execute a sequence of actions, such as reasoning steps or tool calls, in response to a user prompt. To evaluate the success of their trajectories, researchers have developed verifiers, such as LLM judges and process-reward models, to score the quality of each action in an agent's trajectory. Although these heuristic scores can be informative, there are no guarantees of correctness when used to decide whether an agent will yield a successful output. Here, we introduce e-valuator, a method to convert any black-box verifier score into a decision rule with provable control of false alarm rates. We frame the problem of distinguishing successful trajectories (that is, a sequence of actions that will lead to a correct response to the user's prompt) and unsuccessful trajectories as a sequential hypothesis testing problem. E-valuator builds on tools from e-processes to develop a sequential hypothesis test that remains statistically valid at every step of an agent's trajectory, enabling online monitoring of agents over arbitrarily long sequences of actions. Empirically, we demonstrate that e-valuator provides greater statistical power and better false alarm rate control than other strategies across six datasets and three agents. We additionally show that e-valuator can be used for to quickly terminate problematic trajectories and save tokens. Together, e-valuator provides a lightweight, model-agnostic framework that converts verifier heuristics into decisions rules with statistical guarantees, enabling the deployment of more reliable agentic systems.
format Preprint
id arxiv_https___arxiv_org_abs_2512_03109
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing
Sadhuka, Shuvom
Prinster, Drew
Fannjiang, Clara
Scalia, Gabriele
Berger, Bonnie
Regev, Aviv
Wang, Hanchen
Machine Learning
Artificial Intelligence
Applications
Agentic AI systems execute a sequence of actions, such as reasoning steps or tool calls, in response to a user prompt. To evaluate the success of their trajectories, researchers have developed verifiers, such as LLM judges and process-reward models, to score the quality of each action in an agent's trajectory. Although these heuristic scores can be informative, there are no guarantees of correctness when used to decide whether an agent will yield a successful output. Here, we introduce e-valuator, a method to convert any black-box verifier score into a decision rule with provable control of false alarm rates. We frame the problem of distinguishing successful trajectories (that is, a sequence of actions that will lead to a correct response to the user's prompt) and unsuccessful trajectories as a sequential hypothesis testing problem. E-valuator builds on tools from e-processes to develop a sequential hypothesis test that remains statistically valid at every step of an agent's trajectory, enabling online monitoring of agents over arbitrarily long sequences of actions. Empirically, we demonstrate that e-valuator provides greater statistical power and better false alarm rate control than other strategies across six datasets and three agents. We additionally show that e-valuator can be used for to quickly terminate problematic trajectories and save tokens. Together, e-valuator provides a lightweight, model-agnostic framework that converts verifier heuristics into decisions rules with statistical guarantees, enabling the deployment of more reliable agentic systems.
title E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing
topic Machine Learning
Artificial Intelligence
Applications
url https://arxiv.org/abs/2512.03109