APEX-Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vidgen, Bertie, Mann, Austin, Fennelly, Abby, Stanly, John Wright, Rothman, Lucas, Burstein, Marco, Benchek, Julien, Ostrofsky, David, Ravichandran, Anirudh, Sur, Debnil, Venugopal, Neel, Hsia, Alannah, Robinson, Isaac, Huang, Calix, Varones, Olivia, Khan, Daniyal, Haines, Michael, Bridges, Austin, Boyle, Jesse, Twist, Koby, Richards, Zach, Mahapatra, Chirag, Foody, Brendan, Nitski, Osvald
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914343159857152
author Vidgen, Bertie
Mann, Austin
Fennelly, Abby
Stanly, John Wright
Rothman, Lucas
Burstein, Marco
Benchek, Julien
Ostrofsky, David
Ravichandran, Anirudh
Sur, Debnil
Venugopal, Neel
Hsia, Alannah
Robinson, Isaac
Huang, Calix
Varones, Olivia
Khan, Daniyal
Haines, Michael
Bridges, Austin
Boyle, Jesse
Twist, Koby
Richards, Zach
Mahapatra, Chirag
Foody, Brendan
Nitski, Osvald
author_facet Vidgen, Bertie
Mann, Austin
Fennelly, Abby
Stanly, John Wright
Rothman, Lucas
Burstein, Marco
Benchek, Julien
Ostrofsky, David
Ravichandran, Anirudh
Sur, Debnil
Venugopal, Neel
Hsia, Alannah
Robinson, Isaac
Huang, Calix
Varones, Olivia
Khan, Daniyal
Haines, Michael
Bridges, Austin
Boyle, Jesse
Twist, Koby
Richards, Zach
Mahapatra, Chirag
Foody, Brendan
Nitski, Osvald
contents We introduce the AI Productivity Index for Agents (APEX-Agents), a benchmark for assessing whether AI agents can execute long-horizon, cross-application tasks created by investment banking analysts, management consultants, and corporate lawyers. APEX-Agents requires agents to navigate realistic work environments with files and tools. We test eight agents for the leaderboard using Pass@1. Gemini 3 Flash (Thinking=High) achieves the highest score of 24.0%, followed by GPT-5.2 (Thinking=High), Claude Opus 4.5 (Thinking=High), and Gemini 3 Pro (Thinking=High). We open source the APEX-Agents benchmark (n=480) with all prompts, rubrics, gold outputs, files, and metadata. We also open source Archipelago, our infrastructure for agent execution and evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14242
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle APEX-Agents
Vidgen, Bertie
Mann, Austin
Fennelly, Abby
Stanly, John Wright
Rothman, Lucas
Burstein, Marco
Benchek, Julien
Ostrofsky, David
Ravichandran, Anirudh
Sur, Debnil
Venugopal, Neel
Hsia, Alannah
Robinson, Isaac
Huang, Calix
Varones, Olivia
Khan, Daniyal
Haines, Michael
Bridges, Austin
Boyle, Jesse
Twist, Koby
Richards, Zach
Mahapatra, Chirag
Foody, Brendan
Nitski, Osvald
Computation and Language
Artificial Intelligence
Machine Learning
We introduce the AI Productivity Index for Agents (APEX-Agents), a benchmark for assessing whether AI agents can execute long-horizon, cross-application tasks created by investment banking analysts, management consultants, and corporate lawyers. APEX-Agents requires agents to navigate realistic work environments with files and tools. We test eight agents for the leaderboard using Pass@1. Gemini 3 Flash (Thinking=High) achieves the highest score of 24.0%, followed by GPT-5.2 (Thinking=High), Claude Opus 4.5 (Thinking=High), and Gemini 3 Pro (Thinking=High). We open source the APEX-Agents benchmark (n=480) with all prompts, rubrics, gold outputs, files, and metadata. We also open source Archipelago, our infrastructure for agent execution and evaluation.
title APEX-Agents
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2601.14242