ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xiangyi, Choe, Kyoung Whan, Liu, Yimin, Chen, Xiaokun, Tao, Chujun, You, Bingran, Chen, Wenbo, Di, Zonglin, Sun, Jiankai, Zheng, Shenghan, Bao, Jiajun, Wang, Yuanli, Yan, Weixiang, Li, Yiyuan, Lee, Han-chung |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
by: Li, Xiangyi, et al.
Published: (2026)
by: Li, Xiangyi, et al.
Published: (2026)
Massively Multiagent Minigames for Training Generalist Agents
by: Choe, Kyoung Whan, et al.
Published: (2024)
by: Choe, Kyoung Whan, et al.
Published: (2024)
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
by: Tang, Zirui, et al.
Published: (2026)
by: Tang, Zirui, et al.
Published: (2026)
Information Retrieval in Multimedia Sources in an Electronic Age.
by: Li, Tze-chung
Published: (1988)
by: Li, Tze-chung
Published: (1988)
Library Automation in the Republic of China: Practical Aspects and Perspectives.
by: Li, Tze-chung
Published: (1991)
by: Li, Tze-chung
Published: (1991)
The Future of American Library Education--Return to Basics?
by: Li, Tze-chung
Published: (1985)
by: Li, Tze-chung
Published: (1985)
BloClaw: An Omniscient, Multi-Modal Agentic Workspace for Next-Generation Scientific Discovery
by: Qin, Yao, et al.
Published: (2026)
by: Qin, Yao, et al.
Published: (2026)
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
by: Liu, Songyang, et al.
Published: (2026)
by: Liu, Songyang, et al.
Published: (2026)
CleanUpBench: Embodied Sweeping and Grasping Benchmark
by: Li, Wenbo, et al.
Published: (2025)
by: Li, Wenbo, et al.
Published: (2025)
MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows
by: Shen, Weixiang, et al.
Published: (2026)
by: Shen, Weixiang, et al.
Published: (2026)
Privacy Preserving Conversion Modeling in Data Clean Room
by: Li, Kungang, et al.
Published: (2025)
by: Li, Kungang, et al.
Published: (2025)
Europe Style Wooden Clock
by: chung_the_artist
Published: (2020)
by: chung_the_artist
Published: (2020)
Orbital-driven emergent transport in altermagnets
by: Choi, Junyeong, et al.
Published: (2026)
by: Choi, Junyeong, et al.
Published: (2026)
WLPCM Approach for Great Lakes Regulation
by: Chen, Xiangyi, et al.
Published: (2025)
by: Chen, Xiangyi, et al.
Published: (2025)
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis
by: Zhang, Qiaohong, et al.
Published: (2026)
by: Zhang, Qiaohong, et al.
Published: (2026)
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex
by: Yang, Zhonghao, et al.
Published: (2026)
by: Yang, Zhonghao, et al.
Published: (2026)
LR^2Bench: Evaluating Long-chain Reflective Reasoning Capabilities of Large Language Models via Constraint Satisfaction Problems
by: Chen, Jianghao, et al.
Published: (2025)
by: Chen, Jianghao, et al.
Published: (2025)
A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)
by: Chen, Tianyu, et al.
Published: (2026)
by: Chen, Tianyu, et al.
Published: (2026)
ClawBench: Can AI Agents Complete Everyday Online Tasks?
by: Zhang, Yuxuan, et al.
Published: (2026)
by: Zhang, Yuxuan, et al.
Published: (2026)
Considerations in Using Computer for Presentation.
by: Lee, Shih-chung
Published: (1997)
by: Lee, Shih-chung
Published: (1997)
Yield More and Feed More: Unraveling the Multi‐Scale Determinants of the Spatial Variation in Plastic‐Mulched Farmland Across China
by: Yingnan Zhang, et al.
Published: (2024)
by: Yingnan Zhang, et al.
Published: (2024)
Trace and Edit Relation Associations in GPT
by: Li, Jiahang, et al.
Published: (2023)
by: Li, Jiahang, et al.
Published: (2023)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
by: Long, Xiang, et al.
Published: (2026)
by: Long, Xiang, et al.
Published: (2026)
ClawWorm: Self-Propagating Attacks Across LLM Agent Ecosystems
by: Zhang, Yihao, et al.
Published: (2026)
by: Zhang, Yihao, et al.
Published: (2026)
ClawSafety: "Safe" LLMs, Unsafe Agents
by: Wei, Bowen, et al.
Published: (2026)
by: Wei, Bowen, et al.
Published: (2026)
Relativistic Spin-Lattice Interaction Compatible with Discrete Translation Symmetry in Solids
by: Kim, Bumseop, et al.
Published: (2025)
by: Kim, Bumseop, et al.
Published: (2025)
ClawEnvKit: Automatic Environment Generation for Claw-Like Agents
by: Li, Xirui, et al.
Published: (2026)
by: Li, Xirui, et al.
Published: (2026)
Diffusion wave phenomena and optimal time decay for incompressible viscoelastic flows
by: Li, Shenghan, et al.
Published: (2025)
by: Li, Shenghan, et al.
Published: (2025)
Envisioning Mobile Data Visualization Libraries for Digital Health
by: Lee, Bongshin, et al.
Published: (2026)
by: Lee, Bongshin, et al.
Published: (2026)
Appendix C: Triple-AI Mathematical Validation Report for the GIGL Transparent Universe Model ( Full Version)
by: yeung, tak chung terence
Published: (2025)
by: yeung, tak chung terence
Published: (2025)
Tomographic Parameter Scan of Critical Phase Inversion in a GIGL-Constrained Discrete Grid Model Independent Computational Observation Report by ChatGPT, Grok, Gemini
by: yeung, tak chung terence
Published: (2026)
by: yeung, tak chung terence
Published: (2026)
Planar doubling nodal solutions to the Yamabe equation with maximal rank
by: Li, Yuanli, et al.
Published: (2026)
by: Li, Yuanli, et al.
Published: (2026)
TLB-VFI: Temporal-Aware Latent Brownian Bridge Diffusion for Video Frame Interpolation
by: Lyu, Zonglin, et al.
Published: (2025)
by: Lyu, Zonglin, et al.
Published: (2025)
QuantClaw: Precision Where It Matters for OpenClaw
by: Zhang, Manyi, et al.
Published: (2026)
by: Zhang, Manyi, et al.
Published: (2026)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
by: Jiang, Zhuohang, et al.
Published: (2025)
by: Jiang, Zhuohang, et al.
Published: (2025)
L2Calib: $SE(3)$-Manifold Reinforcement Learning for Robust Extrinsic Calibration with Degenerate Motion Resilience
by: Li, Baorun, et al.
Published: (2025)
by: Li, Baorun, et al.
Published: (2025)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
by: Ding, Shuangrui, et al.
Published: (2026)
by: Ding, Shuangrui, et al.
Published: (2026)
KGN-Pro: Keypoint-Based Grasp Prediction through Probabilistic 2D-3D Correspondence Learning
by: Chen, Bingran, et al.
Published: (2025)
by: Chen, Bingran, et al.
Published: (2025)
SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems
by: Yang, Zonglin, et al.
Published: (2026)
by: Yang, Zonglin, et al.
Published: (2026)
StreamingClaw Technical Report
by: Chen, Jiawei, et al.
Published: (2026)
by: Chen, Jiawei, et al.
Published: (2026)
Similar Items
-
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
by: Li, Xiangyi, et al.
Published: (2026) -
Massively Multiagent Minigames for Training Generalist Agents
by: Choe, Kyoung Whan, et al.
Published: (2024) -
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
by: Tang, Zirui, et al.
Published: (2026) -
Information Retrieval in Multimedia Sources in an Electronic Age.
by: Li, Tze-chung
Published: (1988) -
Library Automation in the Republic of China: Practical Aspects and Perspectives.
by: Li, Tze-chung
Published: (1991)