ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Xiangyi, Choe, Kyoung Whan, Liu, Yimin, Chen, Xiaokun, Tao, Chujun, You, Bingran, Chen, Wenbo, Di, Zonglin, Sun, Jiankai, Zheng, Shenghan, Bao, Jiajun, Wang, Yuanli, Yan, Weixiang, Li, Yiyuan, Lee, Han-chung
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917391209857024
author Li, Xiangyi
Choe, Kyoung Whan
Liu, Yimin
Chen, Xiaokun
Tao, Chujun
You, Bingran
Chen, Wenbo
Di, Zonglin
Sun, Jiankai
Zheng, Shenghan
Bao, Jiajun
Wang, Yuanli
Yan, Weixiang
Li, Yiyuan
Lee, Han-chung
author_facet Li, Xiangyi
Choe, Kyoung Whan
Liu, Yimin
Chen, Xiaokun
Tao, Chujun
You, Bingran
Chen, Wenbo
Di, Zonglin
Sun, Jiankai
Zheng, Shenghan
Bao, Jiajun
Wang, Yuanli
Yan, Weixiang
Li, Yiyuan
Lee, Han-chung
contents Large language model (LLM) agents are increasingly deployed to automate productivity tasks (e.g., email, scheduling, document management), but evaluating them on live services is risky due to potentially irreversible changes. Existing benchmarks rely on simplified environments and fail to capture realistic, stateful, multi-service workflows. We introduce ClawsBench, a benchmark for evaluating and improving LLM agents in realistic productivity settings. It includes five high-fidelity mock services (Gmail, Slack, Google Calendar, Google Docs, Google Drive) with full state management and deterministic snapshot/restore, along with 44 structured tasks covering single-service, cross-service, and safety-critical scenarios. We decompose agent scaffolding into two independent levers (domain skills that inject API knowledge via progressive disclosure, and a meta prompt that coordinates behavior across services) and vary both to measure their separate and combined effects. Experiments across 6 models, 4 agent harnesses, and 33 conditions show that with full scaffolding, agents achieve task success rates of 39-64% but exhibit unsafe action rates of 7-33%. On OpenClaw, the top five models fall within a 10 percentage-point band on task success (53-63%), with unsafe action rates from 7% to 23% and no consistent ordering between the two metrics. We identify eight recurring patterns of unsafe behavior, including multi-step sandbox escalation and silent contract modification. We release the trajectories and future dataset at https://clawsbench.com.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05172
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
Li, Xiangyi
Choe, Kyoung Whan
Liu, Yimin
Chen, Xiaokun
Tao, Chujun
You, Bingran
Chen, Wenbo
Di, Zonglin
Sun, Jiankai
Zheng, Shenghan
Bao, Jiajun
Wang, Yuanli
Yan, Weixiang
Li, Yiyuan
Lee, Han-chung
Artificial Intelligence
Large language model (LLM) agents are increasingly deployed to automate productivity tasks (e.g., email, scheduling, document management), but evaluating them on live services is risky due to potentially irreversible changes. Existing benchmarks rely on simplified environments and fail to capture realistic, stateful, multi-service workflows. We introduce ClawsBench, a benchmark for evaluating and improving LLM agents in realistic productivity settings. It includes five high-fidelity mock services (Gmail, Slack, Google Calendar, Google Docs, Google Drive) with full state management and deterministic snapshot/restore, along with 44 structured tasks covering single-service, cross-service, and safety-critical scenarios. We decompose agent scaffolding into two independent levers (domain skills that inject API knowledge via progressive disclosure, and a meta prompt that coordinates behavior across services) and vary both to measure their separate and combined effects. Experiments across 6 models, 4 agent harnesses, and 33 conditions show that with full scaffolding, agents achieve task success rates of 39-64% but exhibit unsafe action rates of 7-33%. On OpenClaw, the top five models fall within a 10 percentage-point band on task success (53-63%), with unsafe action rates from 7% to 23% and no consistent ordering between the two metrics. We identify eight recurring patterns of unsafe behavior, including multi-step sandbox escalation and silent contract modification. We release the trajectories and future dataset at https://clawsbench.com.
title ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
topic Artificial Intelligence
url https://arxiv.org/abs/2604.05172