Is there a half-life for the success rates of AI agents?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Ord, Toby
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912366646525952
author Ord, Toby
author_facet Ord, Toby
contents Building on the recent empirical work of Kwa et al. (2025), I show that within their suite of research-engineering tasks the performance of AI agents on longer-duration tasks can be explained by an extremely simple mathematical model -- a constant rate of failing during each minute a human would take to do the task. This implies an exponentially declining success rate with the length of the task and that each agent could be characterised by its own half-life. This empirical regularity allows us to estimate the success rate for an agent at different task lengths. And the fact that this model is a good fit for the data is suggestive of the underlying causes of failure on longer tasks -- that they involve increasingly large sets of subtasks where failing any one fails the task. Whether this model applies more generally on other suites of tasks is unknown and an important subject for further work.
format Preprint
id arxiv_https___arxiv_org_abs_2505_05115
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Is there a half-life for the success rates of AI agents?
Ord, Toby
Artificial Intelligence
68T42
I.2.8
Building on the recent empirical work of Kwa et al. (2025), I show that within their suite of research-engineering tasks the performance of AI agents on longer-duration tasks can be explained by an extremely simple mathematical model -- a constant rate of failing during each minute a human would take to do the task. This implies an exponentially declining success rate with the length of the task and that each agent could be characterised by its own half-life. This empirical regularity allows us to estimate the success rate for an agent at different task lengths. And the fact that this model is a good fit for the data is suggestive of the underlying causes of failure on longer tasks -- that they involve increasingly large sets of subtasks where failing any one fails the task. Whether this model applies more generally on other suites of tasks is unknown and an important subject for further work.
title Is there a half-life for the success rates of AI agents?
topic Artificial Intelligence
68T42
I.2.8
url https://arxiv.org/abs/2505.05115