JobBench: Aligning Agent Work With Human Will

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Yuetai, Feng, Yichen, Xu, Zhangchen, Ma, Zixian, Zheng, Kaiyuan, Jiang, Fengqing, Sun, Xinghua, Shao, Rulin, Chen, Zichen, Huang, Yue, Han, Xinyang, Lee, Brian, Xu, Kayla, Zeng, Shenglai, Hua, Hang, Zhang, Xiangliang, Alomair, Basel, Krishna, Ranjay, Zettlemoyer, Luke, Koh, Pang Wei, Ramasubramanian, Bhaskar, Niu, Luyao, Yue, Xiang, Poovendran, Radha
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917533381033984
author Li, Yuetai
Feng, Yichen
Xu, Zhangchen
Ma, Zixian
Zheng, Kaiyuan
Jiang, Fengqing
Sun, Xinghua
Shao, Rulin
Chen, Zichen
Huang, Yue
Han, Xinyang
Lee, Brian
Xu, Kayla
Zeng, Shenglai
Hua, Hang
Zhang, Xiangliang
Alomair, Basel
Krishna, Ranjay
Zettlemoyer, Luke
Koh, Pang Wei
Ramasubramanian, Bhaskar
Niu, Luyao
Yue, Xiang
Poovendran, Radha
author_facet Li, Yuetai
Feng, Yichen
Xu, Zhangchen
Ma, Zixian
Zheng, Kaiyuan
Jiang, Fengqing
Sun, Xinghua
Shao, Rulin
Chen, Zichen
Huang, Yue
Han, Xinyang
Lee, Brian
Xu, Kayla
Zeng, Shenglai
Hua, Hang
Zhang, Xiangliang
Alomair, Basel
Krishna, Ranjay
Zettlemoyer, Luke
Koh, Pang Wei
Ramasubramanian, Bhaskar
Niu, Luyao
Yue, Xiang
Poovendran, Radha
contents Current benchmarks for occupational AI agents are scoped primarily by economic values, telling a replacement story. We introduce JobBench, which evaluates AI agents on the workflows that experts identify as high-priority for delegation, empowering humans based on their needs instead of replacing them with GDP value. JobBench covers 130 agentic tasks across 35 occupations. Each task is packaged as a workspace of heterogeneous reference files, requiring the agent to reason through the cluttered information streams of real professional work. Outputs are graded by a fact-anchored chain of rubrics, averaging 35.6 binary criteria per task. We evaluate 36 models; the strongest, Claude Opus~4.7 under Claude Code, reaches only 45.9 %. We hope JobBench shifts the community's target labour-market effect from replacement to enhancement: building agents that do what humans actually want delegated, not only what is most economically valuable.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26329
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle JobBench: Aligning Agent Work With Human Will
Li, Yuetai
Feng, Yichen
Xu, Zhangchen
Ma, Zixian
Zheng, Kaiyuan
Jiang, Fengqing
Sun, Xinghua
Shao, Rulin
Chen, Zichen
Huang, Yue
Han, Xinyang
Lee, Brian
Xu, Kayla
Zeng, Shenglai
Hua, Hang
Zhang, Xiangliang
Alomair, Basel
Krishna, Ranjay
Zettlemoyer, Luke
Koh, Pang Wei
Ramasubramanian, Bhaskar
Niu, Luyao
Yue, Xiang
Poovendran, Radha
Artificial Intelligence
Current benchmarks for occupational AI agents are scoped primarily by economic values, telling a replacement story. We introduce JobBench, which evaluates AI agents on the workflows that experts identify as high-priority for delegation, empowering humans based on their needs instead of replacing them with GDP value. JobBench covers 130 agentic tasks across 35 occupations. Each task is packaged as a workspace of heterogeneous reference files, requiring the agent to reason through the cluttered information streams of real professional work. Outputs are graded by a fact-anchored chain of rubrics, averaging 35.6 binary criteria per task. We evaluate 36 models; the strongest, Claude Opus~4.7 under Claude Code, reaches only 45.9 %. We hope JobBench shifts the community's target labour-market effect from replacement to enhancement: building agents that do what humans actually want delegated, not only what is most economically valuable.
title JobBench: Aligning Agent Work With Human Will
topic Artificial Intelligence
url https://arxiv.org/abs/2605.26329