The AI Productivity Index (APEX)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vidgen, Bertie, Fennelly, Abby, Pinnix, Evan, Benchek, Julien, Khan, Daniyal, Richards, Zach, Bridges, Austin, Huang, Calix, Sahu, Kanishka, Kottamasu, Abhishek, Ma, Bo, Hunsberger, Ben, Robinson, Isaac, Datta, Akul, Mahapatra, Chirag, Barton, Dominic, Sunstein, Cass R., Topol, Eric, Foody, Brendan, Nitski, Osvald
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917150110777344
author Vidgen, Bertie
Fennelly, Abby
Pinnix, Evan
Benchek, Julien
Khan, Daniyal
Richards, Zach
Bridges, Austin
Huang, Calix
Sahu, Kanishka
Kottamasu, Abhishek
Ma, Bo
Hunsberger, Ben
Robinson, Isaac
Datta, Akul
Mahapatra, Chirag
Barton, Dominic
Sunstein, Cass R.
Topol, Eric
Foody, Brendan
Nitski, Osvald
author_facet Vidgen, Bertie
Fennelly, Abby
Pinnix, Evan
Benchek, Julien
Khan, Daniyal
Richards, Zach
Bridges, Austin
Huang, Calix
Sahu, Kanishka
Kottamasu, Abhishek
Ma, Bo
Hunsberger, Ben
Robinson, Isaac
Datta, Akul
Mahapatra, Chirag
Barton, Dominic
Sunstein, Cass R.
Topol, Eric
Foody, Brendan
Nitski, Osvald
contents We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). This technical report details the extensions to APEX-v1, including an increase in the held-out evaluation set from n = 50 to n = 100 cases per job (n = 400 total) and updates to the grading methodology. We present a new leaderboard, where GPT5 (Thinking = High) remains the top performing model with a score of 67.0%. APEX-v1-extended shows that frontier models still have substantial limitations when performing typical professional tasks. To support further research, we are open sourcing n = 25 non-benchmark example cases per role (n = 100 total) along with our evaluation harness.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25721
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The AI Productivity Index (APEX)
Vidgen, Bertie
Fennelly, Abby
Pinnix, Evan
Benchek, Julien
Khan, Daniyal
Richards, Zach
Bridges, Austin
Huang, Calix
Sahu, Kanishka
Kottamasu, Abhishek
Ma, Bo
Hunsberger, Ben
Robinson, Isaac
Datta, Akul
Mahapatra, Chirag
Barton, Dominic
Sunstein, Cass R.
Topol, Eric
Foody, Brendan
Nitski, Osvald
General Economics
Economics
Artificial Intelligence
Computation and Language
Human-Computer Interaction
We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). This technical report details the extensions to APEX-v1, including an increase in the held-out evaluation set from n = 50 to n = 100 cases per job (n = 400 total) and updates to the grading methodology. We present a new leaderboard, where GPT5 (Thinking = High) remains the top performing model with a score of 67.0%. APEX-v1-extended shows that frontier models still have substantial limitations when performing typical professional tasks. To support further research, we are open sourcing n = 25 non-benchmark example cases per role (n = 100 total) along with our evaluation harness.
title The AI Productivity Index (APEX)
topic General Economics
Economics
Artificial Intelligence
Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2509.25721