Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Chenxin, Tang, Zhengyang, Huang, Mingxin, Lin, Yunlong, Huang, Shijue, Liu, Shengyuan, Ye, Bowen, Li, Rang, Li, Lei, Wang, Benyou, Yuan, Yixuan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
by: Meng, Fanqing, et al.
Published: (2026)
by: Meng, Fanqing, et al.
Published: (2026)
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
by: Ye, Bowen, et al.
Published: (2026)
by: Ye, Bowen, et al.
Published: (2026)
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
by: Long, Xiang, et al.
Published: (2026)
by: Long, Xiang, et al.
Published: (2026)
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
by: Zhao, Songwen, et al.
Published: (2025)
by: Zhao, Songwen, et al.
Published: (2025)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
by: Liu, Zhou, et al.
Published: (2025)
by: Liu, Zhou, et al.
Published: (2025)
StatsClaw: An AI-Collaborative Workflow for Statistical Software Development
by: Qin, Tianzhu, et al.
Published: (2026)
by: Qin, Tianzhu, et al.
Published: (2026)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
SecVulEval: Benchmarking LLMs for Real-World C/C++ Vulnerability Detection
by: Ahmed, Md Basim Uddin, et al.
Published: (2025)
by: Ahmed, Md Basim Uddin, et al.
Published: (2025)
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex
by: Yang, Zhonghao, et al.
Published: (2026)
by: Yang, Zhonghao, et al.
Published: (2026)
ClarEval: A Benchmark for Evaluating Clarification Skills of Code Agents under Ambiguous Instructions
by: Li, Jialin, et al.
Published: (2026)
by: Li, Jialin, et al.
Published: (2026)
TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark
by: Jain, Kush, et al.
Published: (2024)
by: Jain, Kush, et al.
Published: (2024)
World of Workflows: A Benchmark for Bringing World Models to Enterprise Systems
by: Gupta, Lakshya, et al.
Published: (2026)
by: Gupta, Lakshya, et al.
Published: (2026)
DependEval: Benchmarking LLMs for Repository Dependency Understanding
by: Du, Junjia, et al.
Published: (2025)
by: Du, Junjia, et al.
Published: (2025)
SWE-bench Goes Live!
by: Zhang, Linghao, et al.
Published: (2025)
by: Zhang, Linghao, et al.
Published: (2025)
Outrunning LLM Cutoffs: A Live Kernel Crash Resolution Benchmark for All
by: Huang, Chenxi, et al.
Published: (2026)
by: Huang, Chenxi, et al.
Published: (2026)
Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?
by: Xia, Chunqiu Steven, et al.
Published: (2025)
by: Xia, Chunqiu Steven, et al.
Published: (2025)
LiveFMBench: Unveiling the Power and Limits of Agentic Workflows in Specification Generation
by: Xu, Dong, et al.
Published: (2026)
by: Xu, Dong, et al.
Published: (2026)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
A Benchmark for Language Models in Real-World System Building
by: Jin, Weilin, et al.
Published: (2026)
by: Jin, Weilin, et al.
Published: (2026)
RepoMasterEval: Evaluating Code Completion via Real-World Repositories
by: Wu, Qinyun, et al.
Published: (2024)
by: Wu, Qinyun, et al.
Published: (2024)
ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation
by: Liu, Kaiyuan, et al.
Published: (2025)
by: Liu, Kaiyuan, et al.
Published: (2025)
CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges
by: Zhang, Kechi, et al.
Published: (2024)
by: Zhang, Kechi, et al.
Published: (2024)
EvoClaw: Evaluating AI Agents on Continuous Software Evolution
by: Deng, Gangda, et al.
Published: (2026)
by: Deng, Gangda, et al.
Published: (2026)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
by: Ni, Ziyi, et al.
Published: (2025)
by: Ni, Ziyi, et al.
Published: (2025)
RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
CompileAgent: Automated Real-World Repo-Level Compilation with Tool-Integrated LLM-based Agent System
by: Hu, Li, et al.
Published: (2025)
by: Hu, Li, et al.
Published: (2025)
HumanEvalComm: Benchmarking the Communication Competence of Code Generation for LLMs and LLM Agent
by: Wu, Jie JW, et al.
Published: (2024)
by: Wu, Jie JW, et al.
Published: (2024)
An Empirical Study of Proactive Coding Assistants in Real-World Software Development
by: Li, Lehui, et al.
Published: (2026)
by: Li, Lehui, et al.
Published: (2026)
Beyond the YAML File: Understanding Real-World GitHub Actions Workflow Adoption
by: Khatami, Ali, et al.
Published: (2026)
by: Khatami, Ali, et al.
Published: (2026)
SELF-REDRAFT: Eliciting Intrinsic Exploration-Exploitation Balance in Test-Time Scaling for Code Generation
by: Chen, Yixiang, et al.
Published: (2025)
by: Chen, Yixiang, et al.
Published: (2025)
Benchmarking LLMs for Unit Test Generation from Real-World Functions
by: Huang, Dong, et al.
Published: (2025)
by: Huang, Dong, et al.
Published: (2025)
SolContractEval: A Benchmark for Evaluating Contract-Level Solidity Code Generation
by: Ye, Zhifan, et al.
Published: (2025)
by: Ye, Zhifan, et al.
Published: (2025)
From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
REVERE: Reflective Evolving Research Engineer for Scientific Workflows
by: Gangireddi, Balaji Dinesh, et al.
Published: (2026)
by: Gangireddi, Balaji Dinesh, et al.
Published: (2026)
How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions
by: Tang, Ningzhi, et al.
Published: (2026)
by: Tang, Ningzhi, et al.
Published: (2026)
CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
by: Luo, Hanjun, et al.
Published: (2025)
by: Luo, Hanjun, et al.
Published: (2025)
CoderEval: A Benchmark of Pragmatic Code Generation with Generative Pre-trained Models
by: Yu, Hao, et al.
Published: (2023)
by: Yu, Hao, et al.
Published: (2023)
Benchmarking and Studying the LLM-based Agent System in End-to-End Software Development
by: Zeng, Zhengran, et al.
Published: (2025)
by: Zeng, Zhengran, et al.
Published: (2025)
Towards Living Software Architecture Diagrams
by: Correia, Filipe F., et al.
Published: (2024)
by: Correia, Filipe F., et al.
Published: (2024)
Similar Items
-
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
by: Meng, Fanqing, et al.
Published: (2026) -
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
by: Ye, Bowen, et al.
Published: (2026) -
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
by: Li, Jia, et al.
Published: (2024) -
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
by: Long, Xiang, et al.
Published: (2026) -
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
by: Zhao, Songwen, et al.
Published: (2025)