Table of Contents: :: Library Catalog

Saved in:

Bibliographic Details
Main Authors:	Li, Yuetai, Feng, Yichen, Xu, Zhangchen, Ma, Zixian, Zheng, Kaiyuan, Jiang, Fengqing, Sun, Xinghua, Shao, Rulin, Chen, Zichen, Huang, Yue, Han, Xinyang, Lee, Brian, Xu, Kayla, Zeng, Shenglai, Hua, Hang, Zhang, Xiangliang, Alomair, Basel, Krishna, Ranjay, Zettlemoyer, Luke, Koh, Pang Wei, Ramasubramanian, Bhaskar, Niu, Luyao, Yue, Xiang, Poovendran, Radha
Format:	Preprint
Published:	2026
Subjects:	Artificial Intelligence
Online Access:	https://arxiv.org/abs/2605.26329
Tags:	Add Tag No Tags, Be the first to tag this record!

Table of Contents:

Current benchmarks for occupational AI agents are scoped primarily by economic values, telling a replacement story. We introduce JobBench, which evaluates AI agents on the workflows that experts identify as high-priority for delegation, empowering humans based on their needs instead of replacing them with GDP value. JobBench covers 130 agentic tasks across 35 occupations. Each task is packaged as a workspace of heterogeneous reference files, requiring the agent to reason through the cluttered information streams of real professional work. Outputs are graded by a fact-anchored chain of rubrics, averaging 35.6 binary criteria per task. We evaluate 36 models; the strongest, Claude Opus~4.7 under Claude Code, reaches only 45.9 %. We hope JobBench shifts the community's target labour-market effect from replacement to enhancement: building agents that do what humans actually want delegated, not only what is most economically valuable.

Similar Items