Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Chenxin, Tang, Zhengyang, Huang, Mingxin, Lin, Yunlong, Huang, Shijue, Liu, Shengyuan, Ye, Bowen, Li, Rang, Li, Lei, Wang, Benyou, Yuan, Yixuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914523391197184
author Li, Chenxin
Tang, Zhengyang
Huang, Mingxin
Lin, Yunlong
Huang, Shijue
Liu, Shengyuan
Ye, Bowen
Li, Rang
Li, Lei
Wang, Benyou
Yuan, Yixuan
author_facet Li, Chenxin
Tang, Zhengyang
Huang, Mingxin
Lin, Yunlong
Huang, Shijue
Liu, Shengyuan
Ye, Bowen
Li, Rang
Li, Lei
Wang, Benyou
Yuan, Yixuan
contents LLM agents are expected to complete end-to-end units of work across software tools, business services, and local workspaces. Yet many agent benchmarks freeze a curated task set at release time and grade mainly the final response, making it difficult to evaluate agents against evolving workflow demand or verify whether a task was executed. We introduce Claw-Eval-Live, a live benchmark for workflow agents that separates a refreshable signal layer, updated across releases from public workflow-demand signals, from a reproducible, time-stamped release snapshot. Each release is constructed from public workflow-demand signals, with ClawHub Top-500 skills used in the current release, and materialized as controlled tasks with fixed fixtures, services, workspaces, and graders. For grading, Claw-Eval-Live records execution traces, audit logs, service state, and post-run workspace artifacts, using deterministic checks when evidence is sufficient and structured LLM judging only for semantic dimensions. The release contains 105 tasks spanning controlled business services and local workspace repair, and evaluates 13 frontier models under a shared public pass rule. Experiments reveal that reliable workflow automation remains far from solved: the leading model passes only 66.7% of tasks and no model reaches 70%. Failures are structured by task family and execution surface, with HR, management, and multi-system business workflows as persistent bottlenecks and local workspace repair comparatively easier but unsaturated. Leaderboard rank alone is insufficient because models with similar pass rates can diverge in overall completion, and task-level discrimination concentrates in a middle band of tasks. Claw-Eval-Live suggests that workflow-agent evaluation should be grounded twice, in fresh external demand and in verifiable agent action.
format Preprint
id arxiv_https___arxiv_org_abs_2604_28139
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
Li, Chenxin
Tang, Zhengyang
Huang, Mingxin
Lin, Yunlong
Huang, Shijue
Liu, Shengyuan
Ye, Bowen
Li, Rang
Li, Lei
Wang, Benyou
Yuan, Yixuan
Software Engineering
Artificial Intelligence
LLM agents are expected to complete end-to-end units of work across software tools, business services, and local workspaces. Yet many agent benchmarks freeze a curated task set at release time and grade mainly the final response, making it difficult to evaluate agents against evolving workflow demand or verify whether a task was executed. We introduce Claw-Eval-Live, a live benchmark for workflow agents that separates a refreshable signal layer, updated across releases from public workflow-demand signals, from a reproducible, time-stamped release snapshot. Each release is constructed from public workflow-demand signals, with ClawHub Top-500 skills used in the current release, and materialized as controlled tasks with fixed fixtures, services, workspaces, and graders. For grading, Claw-Eval-Live records execution traces, audit logs, service state, and post-run workspace artifacts, using deterministic checks when evidence is sufficient and structured LLM judging only for semantic dimensions. The release contains 105 tasks spanning controlled business services and local workspace repair, and evaluates 13 frontier models under a shared public pass rule. Experiments reveal that reliable workflow automation remains far from solved: the leading model passes only 66.7% of tasks and no model reaches 70%. Failures are structured by task family and execution surface, with HR, management, and multi-system business workflows as persistent bottlenecks and local workspace repair comparatively easier but unsaturated. Leaderboard rank alone is insufficient because models with similar pass rates can diverge in overall completion, and task-level discrimination concentrates in a middle band of tasks. Claw-Eval-Live suggests that workflow-agent evaluation should be grounded twice, in fresh external demand and in verifiable agent action.
title Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2604.28139