Saved in:
| Main Authors: | Chen, Tianyu, Hu, Chujia, Gao, Ge, Liu, Dongrui, Hu, Xia, Wang, Wenjie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.03255 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)
by: Chen, Tianyu, et al.
Published: (2026)
by: Chen, Tianyu, et al.
Published: (2026)
UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios
by: Luo, Haotian, et al.
Published: (2025)
by: Luo, Haotian, et al.
Published: (2025)
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
by: Shen, Yuanzhe, et al.
Published: (2026)
by: Shen, Yuanzhe, et al.
Published: (2026)
CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios
by: Cai, Muzhen, et al.
Published: (2025)
by: Cai, Muzhen, et al.
Published: (2025)
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
by: Yang, Jingyi, et al.
Published: (2025)
by: Yang, Jingyi, et al.
Published: (2025)
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios
by: Qiu, Lu, et al.
Published: (2024)
by: Qiu, Lu, et al.
Published: (2024)
IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
by: Lu, Xiaoya, et al.
Published: (2025)
by: Lu, Xiaoya, et al.
Published: (2025)
OS-Marathon: Benchmarking Computer-Use Agents on Long-Horizon Repetitive Tasks
by: Wu, Jing, et al.
Published: (2026)
by: Wu, Jing, et al.
Published: (2026)
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
by: Jin, Chang, et al.
Published: (2026)
by: Jin, Chang, et al.
Published: (2026)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
by: Ding, Shuangrui, et al.
Published: (2026)
by: Ding, Shuangrui, et al.
Published: (2026)
MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios
by: Song, Zhiheng, et al.
Published: (2026)
by: Song, Zhiheng, et al.
Published: (2026)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
by: Orlanski, Gabriel, et al.
Published: (2026)
by: Orlanski, Gabriel, et al.
Published: (2026)
Adversarial Artifact Detection in EEG-Based Brain-Computer Interfaces
by: Chen, Xiaoqing, et al.
Published: (2022)
by: Chen, Xiaoqing, et al.
Published: (2022)
HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds
by: Anokhin, Petr, et al.
Published: (2025)
by: Anokhin, Petr, et al.
Published: (2025)
LongGenBench: Long-context Generation Benchmark
by: Liu, Xiang, et al.
Published: (2024)
by: Liu, Xiang, et al.
Published: (2024)
ELHPlan: Efficient Long-Horizon Task Planning for Multi-Agent Collaboration
by: Ling, Shaobin, et al.
Published: (2025)
by: Ling, Shaobin, et al.
Published: (2025)
The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
by: Ding, Xuwei, et al.
Published: (2026)
by: Ding, Xuwei, et al.
Published: (2026)
TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety
by: Hong, Zhepei, et al.
Published: (2026)
by: Hong, Zhepei, et al.
Published: (2026)
GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning
by: Cheng, Xiang, et al.
Published: (2026)
by: Cheng, Xiang, et al.
Published: (2026)
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Synthetic Computers at Scale for Long-Horizon Productivity Simulation
by: Ge, Tao, et al.
Published: (2026)
by: Ge, Tao, et al.
Published: (2026)
Enhancing Incipient Fault Detection for Interface Converter Sensors through Signal Correlation Analysis
by: Chujia Guo, et al.
Published: (2024)
by: Chujia Guo, et al.
Published: (2024)
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
by: Song, Yuanyi, et al.
Published: (2025)
by: Song, Yuanyi, et al.
Published: (2025)
LifeBench: A Benchmark for Long-Horizon Multi-Source Memory
by: Cheng, Zihao, et al.
Published: (2026)
by: Cheng, Zihao, et al.
Published: (2026)
AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents
by: Hu, Lingxiang, et al.
Published: (2026)
by: Hu, Lingxiang, et al.
Published: (2026)
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
by: Gao, Yuxuan, et al.
Published: (2026)
by: Gao, Yuxuan, et al.
Published: (2026)
NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
by: Ding, Jingzhe, et al.
Published: (2025)
by: Ding, Jingzhe, et al.
Published: (2025)
ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario
by: Zhong, Lucen, et al.
Published: (2025)
by: Zhong, Lucen, et al.
Published: (2025)
TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems
by: Xu, Chen, et al.
Published: (2026)
by: Xu, Chen, et al.
Published: (2026)
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex
by: Yang, Zhonghao, et al.
Published: (2026)
by: Yang, Zhonghao, et al.
Published: (2026)
Benign Samples Matter! Fine-tuning On Outlier Benign Samples Severely Breaks Safety
by: Guan, Zihan, et al.
Published: (2025)
by: Guan, Zihan, et al.
Published: (2025)
EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
by: Chen, Yi, et al.
Published: (2023)
by: Chen, Yi, et al.
Published: (2023)
HorizonBench: Long-Horizon Personalization with Evolving Preferences
by: Li, Shuyue Stella, et al.
Published: (2026)
by: Li, Shuyue Stella, et al.
Published: (2026)
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
by: Chen, Ziyang, et al.
Published: (2026)
by: Chen, Ziyang, et al.
Published: (2026)
Empowering LLMs with Parameterized Skills for Adversarial Long-Horizon Planning
by: Cui, Sijia, et al.
Published: (2025)
by: Cui, Sijia, et al.
Published: (2025)
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents
by: Men, Tianyi, et al.
Published: (2025)
by: Men, Tianyi, et al.
Published: (2025)
DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation
by: Hu, Zhenyu, et al.
Published: (2026)
by: Hu, Zhenyu, et al.
Published: (2026)
KellyBench: A Benchmark for Long-Horizon Sequential Decision Making
by: Grady, Thomas, et al.
Published: (2026)
by: Grady, Thomas, et al.
Published: (2026)
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
by: Erdogan, Lutfi Eren, et al.
Published: (2025)
by: Erdogan, Lutfi Eren, et al.
Published: (2025)
VLSBench: Unveiling Visual Leakage in Multimodal Safety
by: Hu, Xuhao, et al.
Published: (2024)
by: Hu, Xuhao, et al.
Published: (2024)
Similar Items
-
A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)
by: Chen, Tianyu, et al.
Published: (2026) -
UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios
by: Luo, Haotian, et al.
Published: (2025) -
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
by: Shen, Yuanzhe, et al.
Published: (2026) -
CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios
by: Cai, Muzhen, et al.
Published: (2025) -
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
by: Yang, Jingyi, et al.
Published: (2025)