The AI Productivity Index (APEX)
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917150110777344 |
|---|---|
| author | Vidgen, Bertie Fennelly, Abby Pinnix, Evan Benchek, Julien Khan, Daniyal Richards, Zach Bridges, Austin Huang, Calix Sahu, Kanishka Kottamasu, Abhishek Ma, Bo Hunsberger, Ben Robinson, Isaac Datta, Akul Mahapatra, Chirag Barton, Dominic Sunstein, Cass R. Topol, Eric Foody, Brendan Nitski, Osvald |
| author_facet | Vidgen, Bertie Fennelly, Abby Pinnix, Evan Benchek, Julien Khan, Daniyal Richards, Zach Bridges, Austin Huang, Calix Sahu, Kanishka Kottamasu, Abhishek Ma, Bo Hunsberger, Ben Robinson, Isaac Datta, Akul Mahapatra, Chirag Barton, Dominic Sunstein, Cass R. Topol, Eric Foody, Brendan Nitski, Osvald |
| contents | We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). This technical report details the extensions to APEX-v1, including an increase in the held-out evaluation set from n = 50 to n = 100 cases per job (n = 400 total) and updates to the grading methodology. We present a new leaderboard, where GPT5 (Thinking = High) remains the top performing model with a score of 67.0%. APEX-v1-extended shows that frontier models still have substantial limitations when performing typical professional tasks. To support further research, we are open sourcing n = 25 non-benchmark example cases per role (n = 100 total) along with our evaluation harness. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_25721 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | The AI Productivity Index (APEX) Vidgen, Bertie Fennelly, Abby Pinnix, Evan Benchek, Julien Khan, Daniyal Richards, Zach Bridges, Austin Huang, Calix Sahu, Kanishka Kottamasu, Abhishek Ma, Bo Hunsberger, Ben Robinson, Isaac Datta, Akul Mahapatra, Chirag Barton, Dominic Sunstein, Cass R. Topol, Eric Foody, Brendan Nitski, Osvald General Economics Economics Artificial Intelligence Computation and Language Human-Computer Interaction We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). This technical report details the extensions to APEX-v1, including an increase in the held-out evaluation set from n = 50 to n = 100 cases per job (n = 400 total) and updates to the grading methodology. We present a new leaderboard, where GPT5 (Thinking = High) remains the top performing model with a score of 67.0%. APEX-v1-extended shows that frontier models still have substantial limitations when performing typical professional tasks. To support further research, we are open sourcing n = 25 non-benchmark example cases per role (n = 100 total) along with our evaluation harness. |
| title | The AI Productivity Index (APEX) |
| topic | General Economics Economics Artificial Intelligence Computation and Language Human-Computer Interaction |
| url | https://arxiv.org/abs/2509.25721 |