AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Jingwei, Zhu, Jianing, Li, Yuanyi, Liu, Tongliang, HU, Xia, Han, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
Rethinking How to Remember: Beyond Atomic Facts in Lifelong LLM Agent Memory
by: Sun, Jingwei, et al.
Published: (2026)
by: Sun, Jingwei, et al.
Published: (2026)
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
by: Yang, Jingyi, et al.
Published: (2025)
by: Yang, Jingyi, et al.
Published: (2025)
C-World: A Computer Use Agent Environment Creator
by: Xi, Ziqiao, et al.
Published: (2026)
by: Xi, Ziqiao, et al.
Published: (2026)
Automating Agent Hijacking via Structural Template Injection
by: Deng, Xinhao, et al.
Published: (2026)
by: Deng, Xinhao, et al.
Published: (2026)
On the Reliability of Computer Use Agents
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2026)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2026)
WebTrap: Stealthy Mid-Task Hijacking of Browser Agents During Navigation
by: Liu, Zhichao, et al.
Published: (2026)
by: Liu, Zhichao, et al.
Published: (2026)
LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios
by: Chen, Tianyu, et al.
Published: (2026)
by: Chen, Tianyu, et al.
Published: (2026)
ClawArena: Benchmarking AI Agents in Evolving Information Environments
by: Ji, Haonian, et al.
Published: (2026)
by: Ji, Haonian, et al.
Published: (2026)
MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents
by: Yan, Yunhe, et al.
Published: (2025)
by: Yan, Yunhe, et al.
Published: (2025)
Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction
by: Wang, Hongtao, et al.
Published: (2026)
by: Wang, Hongtao, et al.
Published: (2026)
TermiGen: High-Fidelity Environment and Robust Trajectory Synthesis for Terminal Agents
by: Zhu, Kaijie, et al.
Published: (2026)
by: Zhu, Kaijie, et al.
Published: (2026)
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
by: Wang, Bowen, et al.
Published: (2026)
by: Wang, Bowen, et al.
Published: (2026)
UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning
by: Zhang, Jiawei, et al.
Published: (2025)
by: Zhang, Jiawei, et al.
Published: (2025)
Benchmarking Agents in Insurance Underwriting Environments
by: Dsouza, Amanda, et al.
Published: (2026)
by: Dsouza, Amanda, et al.
Published: (2026)
OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents
by: Abhyankar, Reyna, et al.
Published: (2025)
by: Abhyankar, Reyna, et al.
Published: (2025)
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
by: Xie, Tianbao, et al.
Published: (2024)
by: Xie, Tianbao, et al.
Published: (2024)
PERMA: Benchmarking Personalized Memory Agents via Event-Driven Preference and Realistic Task Environments
by: Liu, Shuochen, et al.
Published: (2026)
by: Liu, Shuochen, et al.
Published: (2026)
EntWorld: A Holistic Environment and Benchmark for Verifiable Enterprise GUI Agents
by: Mo, Ying, et al.
Published: (2026)
by: Mo, Ying, et al.
Published: (2026)
AgentOrchestra: Orchestrating Multi-Agent Intelligence with the Tool-Environment-Agent(TEA) Protocol
by: Zhang, Wentao, et al.
Published: (2025)
by: Zhang, Wentao, et al.
Published: (2025)
AgentAuditor: Human-Level Safety and Security Evaluation for LLM Agents
by: Luo, Hanjun, et al.
Published: (2025)
by: Luo, Hanjun, et al.
Published: (2025)
PRO-CUA: Process-Reward Optimization for Computer Use Agents
by: He, Yifei, et al.
Published: (2026)
by: He, Yifei, et al.
Published: (2026)
OODFace: Benchmarking Robustness of Face Recognition under Common Corruptions and Appearance Variations
by: Kang, Caixin, et al.
Published: (2024)
by: Kang, Caixin, et al.
Published: (2024)
WebPII: Benchmarking Visual PII Detection for Computer-Use Agents
by: Zhao, Nathan
Published: (2026)
by: Zhao, Nathan
Published: (2026)
ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
by: Lai, Hanyu, et al.
Published: (2025)
by: Lai, Hanyu, et al.
Published: (2025)
TClone: Low-Latency Forking of Live GUI Environments for Computer-Use Agents
by: Huang, Yutong, et al.
Published: (2026)
by: Huang, Yutong, et al.
Published: (2026)
ColorAgent: Building A Robust, Personalized, and Interactive OS Agent
by: Li, Ning, et al.
Published: (2025)
by: Li, Ning, et al.
Published: (2025)
MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
by: Yang, Xiao, et al.
Published: (2025)
by: Yang, Xiao, et al.
Published: (2025)
OpenOmni: A Collaborative Open Source Tool for Building Future-Ready Multimodal Conversational Agents
by: Sun, Qiang, et al.
Published: (2024)
by: Sun, Qiang, et al.
Published: (2024)
ScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent Training
by: Tu, Dunwei, et al.
Published: (2026)
by: Tu, Dunwei, et al.
Published: (2026)
Grounding Computer Use Agents on Human Demonstrations
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
I Can't Believe It's Corrupt: Evaluating Corruption in Multi-Agent Governance Systems
by: P, Vedanta S, et al.
Published: (2026)
by: P, Vedanta S, et al.
Published: (2026)
Efficient Agent Training for Computer Use
by: He, Yanheng, et al.
Published: (2025)
by: He, Yanheng, et al.
Published: (2025)
Belief Memory: Agent Memory Under Partial Observability
by: Liao, Junfeng, et al.
Published: (2026)
by: Liao, Junfeng, et al.
Published: (2026)
Ready Jurist One: Benchmarking Language Agents for Legal Intelligence in Dynamic Environments
by: Jia, Zheng, et al.
Published: (2025)
by: Jia, Zheng, et al.
Published: (2025)
AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems
by: Zhang, Boxuan, et al.
Published: (2026)
by: Zhang, Boxuan, et al.
Published: (2026)
OSExpert: Computer-Use Agents Learning Professional Skills via Exploration
by: Liu, Jiateng, et al.
Published: (2026)
by: Liu, Jiateng, et al.
Published: (2026)
GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis
by: Yu, Bo, et al.
Published: (2026)
by: Yu, Bo, et al.
Published: (2026)
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
by: Liu, Wenrui, et al.
Published: (2025)
by: Liu, Wenrui, et al.
Published: (2025)
PhoneWorld: Scaling Phone-Use Agent Environments
by: Tang, Zhengyang, et al.
Published: (2026)
by: Tang, Zhengyang, et al.
Published: (2026)
Similar Items
-
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
by: Feng, Yunhao, et al.
Published: (2026) -
Rethinking How to Remember: Beyond Atomic Facts in Lifelong LLM Agent Memory
by: Sun, Jingwei, et al.
Published: (2026) -
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
by: Yang, Jingyi, et al.
Published: (2025) -
C-World: A Computer Use Agent Environment Creator
by: Xi, Ziqiao, et al.
Published: (2026) -
Automating Agent Hijacking via Structural Template Injection
by: Deng, Xinhao, et al.
Published: (2026)