ClawBench: Can AI Agents Complete Everyday Online Tasks?
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yuxuan, Wang, Yubo, Zhu, Yipeng, Du, Penghui, Miao, Junwen, Lu, Xuan, Xu, Wendong, Hao, Yunzhuo, Cai, Songcheng, Wang, Xiaochen, Zhang, Huaisong, Wu, Xian, Lu, Yi, Lei, Minyi, Zou, Kai, Yin, Huifeng, Nie, Ping, Chen, Liang, Jiang, Dongfu, Chen, Wenhu, Allen, Kelsey R. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RewardHarness: Self-Evolving Agentic Post-Training
by: Zhang, Yuxuan, et al.
Published: (2026)
by: Zhang, Yuxuan, et al.
Published: (2026)
Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning
by: Ruan, Chi, et al.
Published: (2025)
by: Ruan, Chi, et al.
Published: (2025)
Watch Before You Answer: Learning from Visually Grounded Post-Training
by: Zhang, Yuxuan, et al.
Published: (2026)
by: Zhang, Yuxuan, et al.
Published: (2026)
ACECODER: Acing Coder RL via Automated Test-Case Synthesis
by: Zeng, Huaye, et al.
Published: (2025)
by: Zeng, Huaye, et al.
Published: (2025)
EvolveCoder: Evolving Test Cases via Adversarial Verification for Code Reinforcement Learning
by: Ruan, Chi, et al.
Published: (2026)
by: Ruan, Chi, et al.
Published: (2026)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
by: Ding, Shuangrui, et al.
Published: (2026)
by: Ding, Shuangrui, et al.
Published: (2026)
Augmenting Black-box LLMs with Medical Textbooks for Biomedical Question Answering
by: Wang, Yubo, et al.
Published: (2023)
by: Wang, Yubo, et al.
Published: (2023)
Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate
by: Wang, Yubo, et al.
Published: (2025)
by: Wang, Yubo, et al.
Published: (2025)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
by: Lu, Yujie, et al.
Published: (2024)
by: Lu, Yujie, et al.
Published: (2024)
Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem
by: Wang, Yubo, et al.
Published: (2025)
by: Wang, Yubo, et al.
Published: (2025)
Stability threshold of Couette flow for Boussinesq equations in $\mathbb{R}^2$
by: Chen, Yubo, et al.
Published: (2025)
by: Chen, Yubo, et al.
Published: (2025)
Quantitative blow-up suppression for the Patlak-Keller-Segel(-Navier-Stokes) system via Couette flow on $\mathbb{R}^2$
by: Chen, Yubo, et al.
Published: (2025)
by: Chen, Yubo, et al.
Published: (2025)
ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations
by: Wang, Yubo, et al.
Published: (2025)
by: Wang, Yubo, et al.
Published: (2025)
MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks
by: Chen, Jiacheng, et al.
Published: (2024)
by: Chen, Jiacheng, et al.
Published: (2024)
General-Reasoner: Advancing LLM Reasoning Across All Domains
by: Ma, Xueguang, et al.
Published: (2025)
by: Ma, Xueguang, et al.
Published: (2025)
QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design
by: Schneider, Benjamin, et al.
Published: (2025)
by: Schneider, Benjamin, et al.
Published: (2025)
VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation
by: Ku, Max, et al.
Published: (2023)
by: Ku, Max, et al.
Published: (2023)
TIGERScore: Towards Building Explainable Metric for All Text Generation Tasks
by: Jiang, Dongfu, et al.
Published: (2023)
by: Jiang, Dongfu, et al.
Published: (2023)
ClawLess: A Security Model of AI Agents
by: Lu, Hongyi, et al.
Published: (2026)
by: Lu, Hongyi, et al.
Published: (2026)
Blow-up suppression for the nematic liquid crystal flow via Couette flow on $\mathbb{R}^2$
by: Chen, Yubo, et al.
Published: (2026)
by: Chen, Yubo, et al.
Published: (2026)
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis
by: Li, Zhuofeng, et al.
Published: (2026)
by: Li, Zhuofeng, et al.
Published: (2026)
Binaphthalene‐Modified C ‐Scorpionate Zinc Catalysts for Ring‐Opening Polymerization of Bio‐Based γ ‐Ketolactones to High‐Molar‐Mass Degradable Polyesters
by: Junwen Xiong, et al.
Published: (2025)
by: Junwen Xiong, et al.
Published: (2025)
Binaphthalene‐Modified C ‐Scorpionate Zinc Catalysts for Ring‐Opening Polymerization of Bio‐Based γ ‐Ketolactones to High‐Molar‐Mass Degradable Polyesters
by: Junwen Xiong, et al.
Published: (2025)
by: Junwen Xiong, et al.
Published: (2025)
VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
by: Jiang, Dongfu, et al.
Published: (2025)
by: Jiang, Dongfu, et al.
Published: (2025)
Why can a hydrophilic polyelectrolyte precipitate and redissolve below the critical micelle concentration of an oppositely-charged surfactant ?
by: Yong, Huaisong
Published: (2024)
by: Yong, Huaisong
Published: (2024)
A Security Analysis of the OpenClaw AI Agent Framework
by: Suwansathit, Surada, et al.
Published: (2026)
by: Suwansathit, Surada, et al.
Published: (2026)
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis
by: Zhang, Qiaohong, et al.
Published: (2026)
by: Zhang, Qiaohong, et al.
Published: (2026)
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
by: Liu, Songyang, et al.
Published: (2026)
by: Liu, Songyang, et al.
Published: (2026)
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
by: Li, Xiangyi, et al.
Published: (2026)
by: Li, Xiangyi, et al.
Published: (2026)
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
by: Cai, Songcheng, et al.
Published: (2026)
by: Cai, Songcheng, et al.
Published: (2026)
GenAI Arena: An Open Evaluation Platform for Generative Models
by: Jiang, Dongfu, et al.
Published: (2024)
by: Jiang, Dongfu, et al.
Published: (2024)
MANTIS: Interleaved Multi-Image Instruction Tuning
by: Jiang, Dongfu, et al.
Published: (2024)
by: Jiang, Dongfu, et al.
Published: (2024)
Convergence analysis of positivity-preserving finite difference scheme for the Flory-Huggins-Cahn-Hilliard equation with dynamical boundary condition
by: Guo, Yunzhuo, et al.
Published: (2025)
by: Guo, Yunzhuo, et al.
Published: (2025)
OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories
by: Liu, Yibing, et al.
Published: (2026)
by: Liu, Yibing, et al.
Published: (2026)
VisCoder2: Building Multi-Language Visualization Coding Agents
by: Ni, Yuansheng, et al.
Published: (2025)
by: Ni, Yuansheng, et al.
Published: (2025)
Energy Consumption Optimization, Response Time Differences and Indicators in Cortical Working Memory Revealed by Nonequilibrium
by: Wang, Xiaochen, et al.
Published: (2024)
by: Wang, Xiaochen, et al.
Published: (2024)
Boosting Few-Shot Segmentation via Instance-Aware Data Augmentation and Local Consensus Guided Cross Attention
by: Guo, Li, et al.
Published: (2024)
by: Guo, Li, et al.
Published: (2024)
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
by: Tang, Fei, et al.
Published: (2026)
by: Tang, Fei, et al.
Published: (2026)
A materials informatics framework based on reduced‐order models for extracting structure–property linkages of additively manufactured continuous fiber‐reinforced polymer composites
by: Yawen Zhang, et al.
Published: (2024)
by: Yawen Zhang, et al.
Published: (2024)
Claw-free bricks that every $b$-invariant edge is solitary
by: Zhang, Yipei, et al.
Published: (2025)
by: Zhang, Yipei, et al.
Published: (2025)
Similar Items
-
RewardHarness: Self-Evolving Agentic Post-Training
by: Zhang, Yuxuan, et al.
Published: (2026) -
Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning
by: Ruan, Chi, et al.
Published: (2025) -
Watch Before You Answer: Learning from Visually Grounded Post-Training
by: Zhang, Yuxuan, et al.
Published: (2026) -
ACECODER: Acing Coder RL via Automated Test-Case Synthesis
by: Zeng, Huaye, et al.
Published: (2025) -
EvolveCoder: Evolving Test Cases via Adversarial Verification for Code Reinforcement Learning
by: Ruan, Chi, et al.
Published: (2026)