LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Tianyu, Hu, Chujia, Gao, Ge, Liu, Dongrui, Hu, Xia, Wang, Wenjie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)
von: Chen, Tianyu, et al.
Veröffentlicht: (2026)
von: Chen, Tianyu, et al.
Veröffentlicht: (2026)
UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios
von: Luo, Haotian, et al.
Veröffentlicht: (2025)
von: Luo, Haotian, et al.
Veröffentlicht: (2025)
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
von: Shen, Yuanzhe, et al.
Veröffentlicht: (2026)
von: Shen, Yuanzhe, et al.
Veröffentlicht: (2026)
CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios
von: Cai, Muzhen, et al.
Veröffentlicht: (2025)
von: Cai, Muzhen, et al.
Veröffentlicht: (2025)
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
von: Yang, Jingyi, et al.
Veröffentlicht: (2025)
von: Yang, Jingyi, et al.
Veröffentlicht: (2025)
OS-Marathon: Benchmarking Computer-Use Agents on Long-Horizon Repetitive Tasks
von: Wu, Jing, et al.
Veröffentlicht: (2026)
von: Wu, Jing, et al.
Veröffentlicht: (2026)
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios
von: Qiu, Lu, et al.
Veröffentlicht: (2024)
von: Qiu, Lu, et al.
Veröffentlicht: (2024)
IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
von: Lu, Xiaoya, et al.
Veröffentlicht: (2025)
von: Lu, Xiaoya, et al.
Veröffentlicht: (2025)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
von: Ding, Shuangrui, et al.
Veröffentlicht: (2026)
von: Ding, Shuangrui, et al.
Veröffentlicht: (2026)
MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios
von: Song, Zhiheng, et al.
Veröffentlicht: (2026)
von: Song, Zhiheng, et al.
Veröffentlicht: (2026)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2026)
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2026)
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
von: Jin, Chang, et al.
Veröffentlicht: (2026)
von: Jin, Chang, et al.
Veröffentlicht: (2026)
HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds
von: Anokhin, Petr, et al.
Veröffentlicht: (2025)
von: Anokhin, Petr, et al.
Veröffentlicht: (2025)
Adversarial Artifact Detection in EEG-Based Brain-Computer Interfaces
von: Chen, Xiaoqing, et al.
Veröffentlicht: (2022)
von: Chen, Xiaoqing, et al.
Veröffentlicht: (2022)
LongGenBench: Long-context Generation Benchmark
von: Liu, Xiang, et al.
Veröffentlicht: (2024)
von: Liu, Xiang, et al.
Veröffentlicht: (2024)
TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety
von: Hong, Zhepei, et al.
Veröffentlicht: (2026)
von: Hong, Zhepei, et al.
Veröffentlicht: (2026)
The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
von: Ding, Xuwei, et al.
Veröffentlicht: (2026)
von: Ding, Xuwei, et al.
Veröffentlicht: (2026)
ELHPlan: Efficient Long-Horizon Task Planning for Multi-Agent Collaboration
von: Ling, Shaobin, et al.
Veröffentlicht: (2025)
von: Ling, Shaobin, et al.
Veröffentlicht: (2025)
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
von: Song, Yuanyi, et al.
Veröffentlicht: (2025)
von: Song, Yuanyi, et al.
Veröffentlicht: (2025)
GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning
von: Cheng, Xiang, et al.
Veröffentlicht: (2026)
von: Cheng, Xiang, et al.
Veröffentlicht: (2026)
LifeBench: A Benchmark for Long-Horizon Multi-Source Memory
von: Cheng, Zihao, et al.
Veröffentlicht: (2026)
von: Cheng, Zihao, et al.
Veröffentlicht: (2026)
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
Synthetic Computers at Scale for Long-Horizon Productivity Simulation
von: Ge, Tao, et al.
Veröffentlicht: (2026)
von: Ge, Tao, et al.
Veröffentlicht: (2026)
AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents
von: Hu, Lingxiang, et al.
Veröffentlicht: (2026)
von: Hu, Lingxiang, et al.
Veröffentlicht: (2026)
NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
von: Ding, Jingzhe, et al.
Veröffentlicht: (2025)
von: Ding, Jingzhe, et al.
Veröffentlicht: (2025)
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario
von: Zhong, Lucen, et al.
Veröffentlicht: (2025)
von: Zhong, Lucen, et al.
Veröffentlicht: (2025)
HorizonBench: Long-Horizon Personalization with Evolving Preferences
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2026)
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2026)
Empowering LLMs with Parameterized Skills for Adversarial Long-Horizon Planning
von: Cui, Sijia, et al.
Veröffentlicht: (2025)
von: Cui, Sijia, et al.
Veröffentlicht: (2025)
Enhancing Incipient Fault Detection for Interface Converter Sensors through Signal Correlation Analysis
von: Chujia Guo, et al.
Veröffentlicht: (2024)
von: Chujia Guo, et al.
Veröffentlicht: (2024)
Benign Samples Matter! Fine-tuning On Outlier Benign Samples Severely Breaks Safety
von: Guan, Zihan, et al.
Veröffentlicht: (2025)
von: Guan, Zihan, et al.
Veröffentlicht: (2025)
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents
von: Men, Tianyi, et al.
Veröffentlicht: (2025)
von: Men, Tianyi, et al.
Veröffentlicht: (2025)
KellyBench: A Benchmark for Long-Horizon Sequential Decision Making
von: Grady, Thomas, et al.
Veröffentlicht: (2026)
von: Grady, Thomas, et al.
Veröffentlicht: (2026)
EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
von: Chen, Yi, et al.
Veröffentlicht: (2023)
von: Chen, Yi, et al.
Veröffentlicht: (2023)
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
von: Chen, Ziyang, et al.
Veröffentlicht: (2026)
von: Chen, Ziyang, et al.
Veröffentlicht: (2026)
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2025)
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2025)
DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation
von: Hu, Zhenyu, et al.
Veröffentlicht: (2026)
von: Hu, Zhenyu, et al.
Veröffentlicht: (2026)
AgentLongBench: A Controllable Long Benchmark For Long-Contexts Agents via Environment Rollouts
von: Fang, Shicheng, et al.
Veröffentlicht: (2026)
von: Fang, Shicheng, et al.
Veröffentlicht: (2026)
DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints
von: Zhang, Yinger, et al.
Veröffentlicht: (2026)
von: Zhang, Yinger, et al.
Veröffentlicht: (2026)
OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
von: Kuntz, Thomas, et al.
Veröffentlicht: (2025)
von: Kuntz, Thomas, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)
von: Chen, Tianyu, et al.
Veröffentlicht: (2026) -
UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios
von: Luo, Haotian, et al.
Veröffentlicht: (2025) -
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
von: Shen, Yuanzhe, et al.
Veröffentlicht: (2026) -
CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios
von: Cai, Muzhen, et al.
Veröffentlicht: (2025) -
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
von: Yang, Jingyi, et al.
Veröffentlicht: (2025)