AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Yunhao, Ding, Yifan, Tan, Yingshui, Ma, Xingjun, Li, Yige, Wu, Yutao, Gao, Yifeng, Zhai, Kun, Guo, Yanming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
BraveGuard: From Open-World Threats to Safer Computer-Use Agents
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
von: Wu, Yutao, et al.
Veröffentlicht: (2025)
von: Wu, Yutao, et al.
Veröffentlicht: (2025)
Identity Lock: Locking API Fine-tuned LLMs With Identity-based Wake Words
von: Su, Hongyu, et al.
Veröffentlicht: (2025)
von: Su, Hongyu, et al.
Veröffentlicht: (2025)
OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
von: Kuntz, Thomas, et al.
Veröffentlicht: (2025)
von: Kuntz, Thomas, et al.
Veröffentlicht: (2025)
ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
von: Wu, Yutao, et al.
Veröffentlicht: (2025)
von: Wu, Yutao, et al.
Veröffentlicht: (2025)
Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs
von: Zheng, Xiang, et al.
Veröffentlicht: (2026)
von: Zheng, Xiang, et al.
Veröffentlicht: (2026)
Human-Guided Harm Recovery for Computer Use Agents
von: Li, Christy, et al.
Veröffentlicht: (2026)
von: Li, Christy, et al.
Veröffentlicht: (2026)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
von: Jones, Jaylen, et al.
Veröffentlicht: (2026)
von: Jones, Jaylen, et al.
Veröffentlicht: (2026)
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
von: Li, Juncheng, et al.
Veröffentlicht: (2025)
von: Li, Juncheng, et al.
Veröffentlicht: (2025)
HazardArena: Evaluating Semantic Safety in Vision-Language-Action Models
von: Chen, Zixing, et al.
Veröffentlicht: (2026)
von: Chen, Zixing, et al.
Veröffentlicht: (2026)
BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
von: Li, Yige, et al.
Veröffentlicht: (2024)
von: Li, Yige, et al.
Veröffentlicht: (2024)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
von: Li, Yige, et al.
Veröffentlicht: (2025)
von: Li, Yige, et al.
Veröffentlicht: (2025)
Internal Safety Collapse in Frontier Large Language Models
von: Wu, Yutao, et al.
Veröffentlicht: (2026)
von: Wu, Yutao, et al.
Veröffentlicht: (2026)
Measuring Harmfulness of Computer-Using Agents
von: Tian, Aaron Xuxiang, et al.
Veröffentlicht: (2025)
von: Tian, Aaron Xuxiang, et al.
Veröffentlicht: (2025)
Position: AI Safety Requires Effective Controllability
von: Li, Yige, et al.
Veröffentlicht: (2026)
von: Li, Yige, et al.
Veröffentlicht: (2026)
Beyond Surface Judgments: Human-Grounded Risk Evaluation of LLM-Generated Disinformation
von: Xu, Zonghuan, et al.
Veröffentlicht: (2026)
von: Xu, Zonghuan, et al.
Veröffentlicht: (2026)
SusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents
von: Guo, Longjie, et al.
Veröffentlicht: (2025)
von: Guo, Longjie, et al.
Veröffentlicht: (2025)
FedEGG: Federated Learning with Explicit Global Guidance
von: Zhai, Kun, et al.
Veröffentlicht: (2024)
von: Zhai, Kun, et al.
Veröffentlicht: (2024)
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
von: Cui, Kaiyuan, et al.
Veröffentlicht: (2026)
von: Cui, Kaiyuan, et al.
Veröffentlicht: (2026)
Mirror: A Multi-Agent System for AI-Assisted Ethics Review
von: Ding, Yifan, et al.
Veröffentlicht: (2026)
von: Ding, Yifan, et al.
Veröffentlicht: (2026)
A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5
von: Ma, Xingjun, et al.
Veröffentlicht: (2026)
von: Ma, Xingjun, et al.
Veröffentlicht: (2026)
AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions
von: Sun, Jingwei, et al.
Veröffentlicht: (2026)
von: Sun, Jingwei, et al.
Veröffentlicht: (2026)
OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents
von: Abhyankar, Reyna, et al.
Veröffentlicht: (2025)
von: Abhyankar, Reyna, et al.
Veröffentlicht: (2025)
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
von: Wei, Jinbiao, et al.
Veröffentlicht: (2026)
von: Wei, Jinbiao, et al.
Veröffentlicht: (2026)
HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
von: Liu, Yuexiao, et al.
Veröffentlicht: (2025)
von: Liu, Yuexiao, et al.
Veröffentlicht: (2025)
The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
von: Ding, Xuwei, et al.
Veröffentlicht: (2026)
von: Ding, Xuwei, et al.
Veröffentlicht: (2026)
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
von: Yang, Jingyi, et al.
Veröffentlicht: (2025)
von: Yang, Jingyi, et al.
Veröffentlicht: (2025)
SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents
von: Ding, Yifeng, et al.
Veröffentlicht: (2026)
von: Ding, Yifeng, et al.
Veröffentlicht: (2026)
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents
von: Shen, Yujiong, et al.
Veröffentlicht: (2026)
von: Shen, Yujiong, et al.
Veröffentlicht: (2026)
OOD-MMSafe: Advancing MLLM Safety from Harmful Intent to Hidden Consequences
von: Wen, Ming, et al.
Veröffentlicht: (2026)
von: Wen, Ming, et al.
Veröffentlicht: (2026)
FedAPT: Federated Adversarial Prompt Tuning for Vision-Language Models
von: Zhai, Kun, et al.
Veröffentlicht: (2025)
von: Zhai, Kun, et al.
Veröffentlicht: (2025)
AnyAttack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models
von: Zhang, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2024)
SIDE: Surrogate Conditional Data Extraction from Diffusion Models
von: Chen, Yunhao, et al.
Veröffentlicht: (2024)
von: Chen, Yunhao, et al.
Veröffentlicht: (2024)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
von: Li, Jiayu, et al.
Veröffentlicht: (2025)
von: Li, Jiayu, et al.
Veröffentlicht: (2025)
ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use
von: Tien, Jeremy, et al.
Veröffentlicht: (2026)
von: Tien, Jeremy, et al.
Veröffentlicht: (2026)
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
von: Huang, Hanxun, et al.
Veröffentlicht: (2025)
von: Huang, Hanxun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents
von: Feng, Yunhao, et al.
Veröffentlicht: (2026) -
BraveGuard: From Open-World Threats to Safer Computer-Use Agents
von: Feng, Yunhao, et al.
Veröffentlicht: (2026) -
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
von: Feng, Yunhao, et al.
Veröffentlicht: (2026) -
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
von: Wu, Yutao, et al.
Veröffentlicht: (2025) -
Identity Lock: Locking API Fine-tuned LLMs With Identity-based Wake Words
von: Su, Hongyu, et al.
Veröffentlicht: (2025)