OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liu, Yibing, Liu, Yangze, Yin, Xiaolong, Wang, Bin, Zhang, Chong, Yin, Hao, Han, Zhongyi |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
TrajAD: Trajectory Anomaly Detection for Trustworthy LLM Agents
par: Liu, Yibing, et autres
Publié: (2026)
par: Liu, Yibing, et autres
Publié: (2026)
ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents
par: Lai, Yuxiang, et autres
Publié: (2026)
par: Lai, Yuxiang, et autres
Publié: (2026)
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis
par: Zhang, Qiaohong, et autres
Publié: (2026)
par: Zhang, Qiaohong, et autres
Publié: (2026)
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
par: Yao, Hongwei, et autres
Publié: (2026)
par: Yao, Hongwei, et autres
Publié: (2026)
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex
par: Yang, Zhonghao, et autres
Publié: (2026)
par: Yang, Zhonghao, et autres
Publié: (2026)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
par: Long, Xiang, et autres
Publié: (2026)
par: Long, Xiang, et autres
Publié: (2026)
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
par: Liu, Wenrui, et autres
Publié: (2025)
par: Liu, Wenrui, et autres
Publié: (2025)
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
par: He, Wei, et autres
Publié: (2025)
par: He, Wei, et autres
Publié: (2025)
A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)
par: Chen, Tianyu, et autres
Publié: (2026)
par: Chen, Tianyu, et autres
Publié: (2026)
From Assistant to Double Agent: Formalizing and Benchmarking Attacks on OpenClaw for Personalized Local AI Agent
par: Wang, Yuhang, et autres
Publié: (2026)
par: Wang, Yuhang, et autres
Publié: (2026)
TimeClaw: A Time-Series AI Agent with Exploratory Execution Learning
par: Liu, Hangchen, et autres
Publié: (2026)
par: Liu, Hangchen, et autres
Publié: (2026)
ClawBench: Can AI Agents Complete Everyday Online Tasks?
par: Zhang, Yuxuan, et autres
Publié: (2026)
par: Zhang, Yuxuan, et autres
Publié: (2026)
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
par: Li, Xiangyi, et autres
Publié: (2026)
par: Li, Xiangyi, et autres
Publié: (2026)
GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis
par: Yu, Bo, et autres
Publié: (2026)
par: Yu, Bo, et autres
Publié: (2026)
HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks
par: Cui, Fan, et autres
Publié: (2026)
par: Cui, Fan, et autres
Publié: (2026)
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
par: Liu, Songyang, et autres
Publié: (2026)
par: Liu, Songyang, et autres
Publié: (2026)
Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw
par: Wang, Zijun, et autres
Publié: (2026)
par: Wang, Zijun, et autres
Publié: (2026)
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
par: Li, Chenxin, et autres
Publié: (2026)
par: Li, Chenxin, et autres
Publié: (2026)
OpenClaw-RL: Train Any Agent Simply by Talking
par: Wang, Yinjie, et autres
Publié: (2026)
par: Wang, Yinjie, et autres
Publié: (2026)
QuantClaw: Precision Where It Matters for OpenClaw
par: Zhang, Manyi, et autres
Publié: (2026)
par: Zhang, Manyi, et autres
Publié: (2026)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
par: Liu, Zhou, et autres
Publié: (2025)
par: Liu, Zhou, et autres
Publié: (2025)
MM-Agent: LLM as Agents for Real-world Mathematical Modeling Problem
par: Liu, Fan, et autres
Publié: (2025)
par: Liu, Fan, et autres
Publié: (2025)
ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents
par: Shen, Haiyang, et autres
Publié: (2024)
par: Shen, Haiyang, et autres
Publié: (2024)
ClawArena: Benchmarking AI Agents in Evolving Information Environments
par: Ji, Haonian, et autres
Publié: (2026)
par: Ji, Haonian, et autres
Publié: (2026)
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
par: Zhang, Zehua, et autres
Publié: (2025)
par: Zhang, Zehua, et autres
Publié: (2025)
Redundant or Necessary? A Benchmark for Detecting Redundant Steps in Agent Trajectories
par: Hu, Minyang, et autres
Publié: (2026)
par: Hu, Minyang, et autres
Publié: (2026)
TrafficClaw: A Generalizable LLM Agent in the Unified Physical Environment for Urban Traffic Control
par: Lai, Siqi, et autres
Publié: (2026)
par: Lai, Siqi, et autres
Publié: (2026)
VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents
par: Chen, Yuhao, et autres
Publié: (2026)
par: Chen, Yuhao, et autres
Publié: (2026)
AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents
par: Hu, Lingxiang, et autres
Publié: (2026)
par: Hu, Lingxiang, et autres
Publié: (2026)
ClawTrap: A MITM-Based Red-Teaming Framework for Real-World OpenClaw Security Evaluation
par: Zhao, Haochen, et autres
Publié: (2026)
par: Zhao, Haochen, et autres
Publié: (2026)
MineAnyBuild: Benchmarking Spatial Planning for Open-world AI Agents
par: Wei, Ziming, et autres
Publié: (2025)
par: Wei, Ziming, et autres
Publié: (2025)
OpenGo: An OpenClaw-Based Robotic Dog with Real-Time Skill Switching
par: Li, Hanbing, et autres
Publié: (2026)
par: Li, Hanbing, et autres
Publié: (2026)
LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services
par: He, Hang, et autres
Publié: (2025)
par: He, Hang, et autres
Publié: (2025)
A Security Analysis of the OpenClaw AI Agent Framework
par: Suwansathit, Surada, et autres
Publié: (2026)
par: Suwansathit, Surada, et autres
Publié: (2026)
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
par: Ye, Bowen, et autres
Publié: (2026)
par: Ye, Bowen, et autres
Publié: (2026)
SGR-Bench: Benchmarking Search Agents on State-Gated Retrieval
par: Li, Ningyuan, et autres
Publié: (2026)
par: Li, Ningyuan, et autres
Publié: (2026)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
par: Wang, Yubang, et autres
Publié: (2026)
par: Wang, Yubang, et autres
Publié: (2026)
ProBench: Benchmarking GUI Agents with Accurate Process Information
par: Yang, Leyang, et autres
Publié: (2025)
par: Yang, Leyang, et autres
Publié: (2025)
LiveAgentBench: Comprehensive Benchmarking of Agentic Systems Across 104 Real-World Challenges
par: Li, Hao, et autres
Publié: (2026)
par: Li, Hao, et autres
Publié: (2026)
TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents
par: Liu, Zhiqiang, et autres
Publié: (2026)
par: Liu, Zhiqiang, et autres
Publié: (2026)
Documents similaires
-
TrajAD: Trajectory Anomaly Detection for Trustworthy LLM Agents
par: Liu, Yibing, et autres
Publié: (2026) -
ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents
par: Lai, Yuxiang, et autres
Publié: (2026) -
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis
par: Zhang, Qiaohong, et autres
Publié: (2026) -
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
par: Yao, Hongwei, et autres
Publié: (2026) -
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex
par: Yang, Zhonghao, et autres
Publié: (2026)