The AI Consumer Index (ACE)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Benchek, Julien, Shetty, Rohit, Hunsberger, Benjamin, Arun, Ajay, Richards, Zach, Foody, Brendan, Nitski, Osvald, Vidgen, Bertie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914189912571904
author Benchek, Julien
Shetty, Rohit
Hunsberger, Benjamin
Arun, Ajay
Richards, Zach
Foody, Brendan
Nitski, Osvald
Vidgen, Bertie
author_facet Benchek, Julien
Shetty, Rohit
Hunsberger, Benjamin
Arun, Ajay
Richards, Zach
Foody, Brendan
Nitski, Osvald
Vidgen, Bertie
contents We introduce the first version of the AI Consumer Index (ACE), a benchmark for assessing whether frontier AI models can perform everyday consumer tasks. ACE contains a hidden heldout set of 400 test cases, split across four consumer activities: shopping, food, gaming, and DIY. We are also open sourcing 80 cases as a devset with a CC-BY license. For the ACE leaderboard we evaluated 10 frontier models (with websearch turned on) using a novel grading methodology that dynamically checks whether relevant parts of the response are grounded in the retrieved web sources. GPT 5 (Thinking = High) is the top-performing model, scoring 56.1%, followed by o3 Pro (Thinking = On) at 55.2% and GPT 5.1 (Thinking = High) at 55.1%. Model scores differ across domains, and in Shopping the top model scores under 50\%. We find that models are prone to hallucinating key information, such as prices. ACE shows a substantial gap between the performance of even the best models and consumers' AI needs.
format Preprint
id arxiv_https___arxiv_org_abs_2512_04921
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The AI Consumer Index (ACE)
Benchek, Julien
Shetty, Rohit
Hunsberger, Benjamin
Arun, Ajay
Richards, Zach
Foody, Brendan
Nitski, Osvald
Vidgen, Bertie
Artificial Intelligence
Computation and Language
Human-Computer Interaction
We introduce the first version of the AI Consumer Index (ACE), a benchmark for assessing whether frontier AI models can perform everyday consumer tasks. ACE contains a hidden heldout set of 400 test cases, split across four consumer activities: shopping, food, gaming, and DIY. We are also open sourcing 80 cases as a devset with a CC-BY license. For the ACE leaderboard we evaluated 10 frontier models (with websearch turned on) using a novel grading methodology that dynamically checks whether relevant parts of the response are grounded in the retrieved web sources. GPT 5 (Thinking = High) is the top-performing model, scoring 56.1%, followed by o3 Pro (Thinking = On) at 55.2% and GPT 5.1 (Thinking = High) at 55.1%. Model scores differ across domains, and in Shopping the top model scores under 50\%. We find that models are prone to hallucinating key information, such as prices. ACE shows a substantial gap between the performance of even the best models and consumers' AI needs.
title The AI Consumer Index (ACE)
topic Artificial Intelligence
Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2512.04921