EVMbench: Evaluating AI Agents on Smart Contract Security

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Justin, Bigger, Andreas, Xu, Xiaohai, Lin, Justin W., Applebaum, Andy, Patwardhan, Tejal, Yukseloglu, Alpin, Watkins, Olivia
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908868019224576
author Wang, Justin
Bigger, Andreas
Xu, Xiaohai
Lin, Justin W.
Applebaum, Andy
Patwardhan, Tejal
Yukseloglu, Alpin
Watkins, Olivia
author_facet Wang, Justin
Bigger, Andreas
Xu, Xiaohai
Lin, Justin W.
Applebaum, Andy
Patwardhan, Tejal
Yukseloglu, Alpin
Watkins, Olivia
contents Smart contracts on public blockchains now manage large amounts of value, and vulnerabilities in these systems can lead to substantial losses. As AI agents become more capable at reading, writing, and running code, it is natural to ask how well they can already navigate this landscape, both in ways that improve security and in ways that might increase risk. We introduce EVMbench, an evaluation that measures the ability of agents to detect, patch, and exploit smart contract vulnerabilities. EVMbench draws on 117 curated vulnerabilities from 40 repositories and, in the most realistic setting, uses programmatic grading based on tests and blockchain state under a local Ethereum execution environment. We evaluate a range of frontier agents and find that they are capable of discovering and exploiting vulnerabilities end-to-end against live blockchain instances. We release code, tasks, and tooling to support continued measurement of these capabilities and future work on security.
format Preprint
id arxiv_https___arxiv_org_abs_2603_04915
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EVMbench: Evaluating AI Agents on Smart Contract Security
Wang, Justin
Bigger, Andreas
Xu, Xiaohai
Lin, Justin W.
Applebaum, Andy
Patwardhan, Tejal
Yukseloglu, Alpin
Watkins, Olivia
Machine Learning
Artificial Intelligence
Cryptography and Security
Smart contracts on public blockchains now manage large amounts of value, and vulnerabilities in these systems can lead to substantial losses. As AI agents become more capable at reading, writing, and running code, it is natural to ask how well they can already navigate this landscape, both in ways that improve security and in ways that might increase risk. We introduce EVMbench, an evaluation that measures the ability of agents to detect, patch, and exploit smart contract vulnerabilities. EVMbench draws on 117 curated vulnerabilities from 40 repositories and, in the most realistic setting, uses programmatic grading based on tests and blockchain state under a local Ethereum execution environment. We evaluate a range of frontier agents and find that they are capable of discovering and exploiting vulnerabilities end-to-end against live blockchain instances. We release code, tasks, and tooling to support continued measurement of these capabilities and future work on security.
title EVMbench: Evaluating AI Agents on Smart Contract Security
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2603.04915