APEX-Agents
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914343159857152 |
|---|---|
| author | Vidgen, Bertie Mann, Austin Fennelly, Abby Stanly, John Wright Rothman, Lucas Burstein, Marco Benchek, Julien Ostrofsky, David Ravichandran, Anirudh Sur, Debnil Venugopal, Neel Hsia, Alannah Robinson, Isaac Huang, Calix Varones, Olivia Khan, Daniyal Haines, Michael Bridges, Austin Boyle, Jesse Twist, Koby Richards, Zach Mahapatra, Chirag Foody, Brendan Nitski, Osvald |
| author_facet | Vidgen, Bertie Mann, Austin Fennelly, Abby Stanly, John Wright Rothman, Lucas Burstein, Marco Benchek, Julien Ostrofsky, David Ravichandran, Anirudh Sur, Debnil Venugopal, Neel Hsia, Alannah Robinson, Isaac Huang, Calix Varones, Olivia Khan, Daniyal Haines, Michael Bridges, Austin Boyle, Jesse Twist, Koby Richards, Zach Mahapatra, Chirag Foody, Brendan Nitski, Osvald |
| contents | We introduce the AI Productivity Index for Agents (APEX-Agents), a benchmark for assessing whether AI agents can execute long-horizon, cross-application tasks created by investment banking analysts, management consultants, and corporate lawyers. APEX-Agents requires agents to navigate realistic work environments with files and tools. We test eight agents for the leaderboard using Pass@1. Gemini 3 Flash (Thinking=High) achieves the highest score of 24.0%, followed by GPT-5.2 (Thinking=High), Claude Opus 4.5 (Thinking=High), and Gemini 3 Pro (Thinking=High). We open source the APEX-Agents benchmark (n=480) with all prompts, rubrics, gold outputs, files, and metadata. We also open source Archipelago, our infrastructure for agent execution and evaluation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_14242 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | APEX-Agents Vidgen, Bertie Mann, Austin Fennelly, Abby Stanly, John Wright Rothman, Lucas Burstein, Marco Benchek, Julien Ostrofsky, David Ravichandran, Anirudh Sur, Debnil Venugopal, Neel Hsia, Alannah Robinson, Isaac Huang, Calix Varones, Olivia Khan, Daniyal Haines, Michael Bridges, Austin Boyle, Jesse Twist, Koby Richards, Zach Mahapatra, Chirag Foody, Brendan Nitski, Osvald Computation and Language Artificial Intelligence Machine Learning We introduce the AI Productivity Index for Agents (APEX-Agents), a benchmark for assessing whether AI agents can execute long-horizon, cross-application tasks created by investment banking analysts, management consultants, and corporate lawyers. APEX-Agents requires agents to navigate realistic work environments with files and tools. We test eight agents for the leaderboard using Pass@1. Gemini 3 Flash (Thinking=High) achieves the highest score of 24.0%, followed by GPT-5.2 (Thinking=High), Claude Opus 4.5 (Thinking=High), and Gemini 3 Pro (Thinking=High). We open source the APEX-Agents benchmark (n=480) with all prompts, rubrics, gold outputs, files, and metadata. We also open source Archipelago, our infrastructure for agent execution and evaluation. |
| title | APEX-Agents |
| topic | Computation and Language Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2601.14242 |