ClawBench: Can AI Agents Complete Everyday Online Tasks?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Yuxuan, Wang, Yubo, Zhu, Yipeng, Du, Penghui, Miao, Junwen, Lu, Xuan, Xu, Wendong, Hao, Yunzhuo, Cai, Songcheng, Wang, Xiaochen, Zhang, Huaisong, Wu, Xian, Lu, Yi, Lei, Minyi, Zou, Kai, Yin, Huifeng, Nie, Ping, Chen, Liang, Jiang, Dongfu, Chen, Wenhu, Allen, Kelsey R.
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908949793472512
author Zhang, Yuxuan
Wang, Yubo
Zhu, Yipeng
Du, Penghui
Miao, Junwen
Lu, Xuan
Xu, Wendong
Hao, Yunzhuo
Cai, Songcheng
Wang, Xiaochen
Zhang, Huaisong
Wu, Xian
Lu, Yi
Lei, Minyi
Zou, Kai
Yin, Huifeng
Nie, Ping
Chen, Liang
Jiang, Dongfu
Chen, Wenhu
Allen, Kelsey R.
author_facet Zhang, Yuxuan
Wang, Yubo
Zhu, Yipeng
Du, Penghui
Miao, Junwen
Lu, Xuan
Xu, Wendong
Hao, Yunzhuo
Cai, Songcheng
Wang, Xiaochen
Zhang, Huaisong
Wu, Xian
Lu, Yi
Lei, Minyi
Zou, Kai
Yin, Huifeng
Nie, Ping
Chen, Liang
Jiang, Dongfu
Chen, Wenhu
Allen, Kelsey R.
contents AI agents may be able to automate your inbox, but can they automate other routine aspects of your life? Everyday online tasks offer a realistic yet unsolved testbed for evaluating the next generation of AI agents. To this end, we introduce ClawBench, an evaluation framework of 153 simple tasks that people need to accomplish regularly in their lives and work, spanning 144 live platforms across 15 categories, from completing purchases and booking appointments to submitting job applications. These tasks require demanding capabilities beyond existing benchmarks, such as obtaining relevant information from user-provided documents, navigating multi-step workflows across diverse platforms, and write-heavy operations like filling in many detailed forms correctly. Unlike existing benchmarks that evaluate agents in offline sandboxes with static pages, ClawBench operates on production websites, preserving the full complexity, dynamic nature, and challenges of real-world web interaction. A lightweight interception layer captures and blocks only the final submission request, ensuring safe evaluation without real-world side effects. Our evaluations of 7 frontier models show that both proprietary and open-source models can complete only a small portion of these tasks. For example, Claude Sonnet 4.6 achieves only 33.3%. Progress on ClawBench brings us closer to AI agents that can function as reliable general-purpose assistants.
format Preprint
id arxiv_https___arxiv_org_abs_2604_08523
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ClawBench: Can AI Agents Complete Everyday Online Tasks?
Zhang, Yuxuan
Wang, Yubo
Zhu, Yipeng
Du, Penghui
Miao, Junwen
Lu, Xuan
Xu, Wendong
Hao, Yunzhuo
Cai, Songcheng
Wang, Xiaochen
Zhang, Huaisong
Wu, Xian
Lu, Yi
Lei, Minyi
Zou, Kai
Yin, Huifeng
Nie, Ping
Chen, Liang
Jiang, Dongfu
Chen, Wenhu
Allen, Kelsey R.
Computation and Language
Artificial Intelligence
AI agents may be able to automate your inbox, but can they automate other routine aspects of your life? Everyday online tasks offer a realistic yet unsolved testbed for evaluating the next generation of AI agents. To this end, we introduce ClawBench, an evaluation framework of 153 simple tasks that people need to accomplish regularly in their lives and work, spanning 144 live platforms across 15 categories, from completing purchases and booking appointments to submitting job applications. These tasks require demanding capabilities beyond existing benchmarks, such as obtaining relevant information from user-provided documents, navigating multi-step workflows across diverse platforms, and write-heavy operations like filling in many detailed forms correctly. Unlike existing benchmarks that evaluate agents in offline sandboxes with static pages, ClawBench operates on production websites, preserving the full complexity, dynamic nature, and challenges of real-world web interaction. A lightweight interception layer captures and blocks only the final submission request, ensuring safe evaluation without real-world side effects. Our evaluations of 7 frontier models show that both proprietary and open-source models can complete only a small portion of these tasks. For example, Claude Sonnet 4.6 achieves only 33.3%. Progress on ClawBench brings us closer to AI agents that can function as reliable general-purpose assistants.
title ClawBench: Can AI Agents Complete Everyday Online Tasks?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.08523