RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Jingyi, Shao, Shuai, Liu, Dongrui, Shao, Jing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents
by: Abhyankar, Reyna, et al.
Published: (2025)
by: Abhyankar, Reyna, et al.
Published: (2025)
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
by: Xie, Tianbao, et al.
Published: (2024)
by: Xie, Tianbao, et al.
Published: (2024)
macOSWorld: A Multilingual Interactive Benchmark for GUI Agents
by: Yang, Pei, et al.
Published: (2025)
by: Yang, Pei, et al.
Published: (2025)
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
by: Shao, Shuai, et al.
Published: (2025)
by: Shao, Shuai, et al.
Published: (2025)
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
by: Guo, Dadi, et al.
Published: (2025)
by: Guo, Dadi, et al.
Published: (2025)
Loop as a Bridge: Can Looped Transformers Truly Link Representation Space and Natural Language Outputs?
by: Chen, Guanxu, et al.
Published: (2026)
by: Chen, Guanxu, et al.
Published: (2026)
ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents
by: Hu, Xuhao, et al.
Published: (2026)
by: Hu, Xuhao, et al.
Published: (2026)
LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios
by: Chen, Tianyu, et al.
Published: (2026)
by: Chen, Tianyu, et al.
Published: (2026)
INFA-Guard: Mitigating Malicious Propagation via Infection-Aware Safeguarding in LLM-Based Multi-Agent Systems
by: Zhou, Yijin, et al.
Published: (2026)
by: Zhou, Yijin, et al.
Published: (2026)
The Tug of War Within: Mitigating the Fairness-Privacy Conflicts in Large Language Models
by: Qian, Chen, et al.
Published: (2024)
by: Qian, Chen, et al.
Published: (2024)
VLSBench: Unveiling Visual Leakage in Multimodal Safety
by: Hu, Xuhao, et al.
Published: (2024)
by: Hu, Xuhao, et al.
Published: (2024)
TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?
by: Yan, Lewen, et al.
Published: (2025)
by: Yan, Lewen, et al.
Published: (2025)
Are Your Agents Upward Deceivers?
by: Guo, Dadi, et al.
Published: (2025)
by: Guo, Dadi, et al.
Published: (2025)
Interpreting Emergent Extreme Events in Multi-Agent Systems
by: Tang, Ling, et al.
Published: (2026)
by: Tang, Ling, et al.
Published: (2026)
ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model Merging
by: Yang, Junyao, et al.
Published: (2026)
by: Yang, Junyao, et al.
Published: (2026)
Attributing Emergence in Million-Agent Systems
by: Tang, Ling, et al.
Published: (2026)
by: Tang, Ling, et al.
Published: (2026)
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
by: Jia, Hongrui, et al.
Published: (2025)
by: Jia, Hongrui, et al.
Published: (2025)
LED-Merging: Mitigating Safety-Utility Conflicts in Model Merging with Location-Election-Disjoint
by: Ma, Qianli, et al.
Published: (2025)
by: Ma, Qianli, et al.
Published: (2025)
Conditional Advantage Estimation for Reinforcement Learning in Large Reasoning Models
by: Chen, Guanxu, et al.
Published: (2025)
by: Chen, Guanxu, et al.
Published: (2025)
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
by: Ren, Qihan, et al.
Published: (2026)
by: Ren, Qihan, et al.
Published: (2026)
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex
by: Yang, Zhonghao, et al.
Published: (2026)
by: Yang, Zhonghao, et al.
Published: (2026)
AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions
by: Sun, Jingwei, et al.
Published: (2026)
by: Sun, Jingwei, et al.
Published: (2026)
Taming Masked Diffusion Language Models via Consistency Trajectory Reinforcement Learning with Fewer Decoding Step
by: Yang, Jingyi, et al.
Published: (2025)
by: Yang, Jingyi, et al.
Published: (2025)
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation
by: Zhou, Tianyi, et al.
Published: (2026)
by: Zhou, Tianyi, et al.
Published: (2026)
Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning
by: Qian, Chen, et al.
Published: (2025)
by: Qian, Chen, et al.
Published: (2025)
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
by: Liu, Zicheng, et al.
Published: (2025)
by: Liu, Zicheng, et al.
Published: (2025)
RiEMann: Near Real-Time SE(3)-Equivariant Robot Manipulation without Point Cloud Segmentation
by: Gao, Chongkai, et al.
Published: (2024)
by: Gao, Chongkai, et al.
Published: (2024)
The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations
by: Zhu, Yubo, et al.
Published: (2025)
by: Zhu, Yubo, et al.
Published: (2025)
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
SPA-Bench: A Comprehensive Benchmark for SmartPhone Agent Evaluation
by: Chen, Jingxuan, et al.
Published: (2024)
by: Chen, Jingxuan, et al.
Published: (2024)
MMSkills: Towards Multimodal Skills for General Visual Agents
by: Zhang, Kangning, et al.
Published: (2026)
by: Zhang, Kangning, et al.
Published: (2026)
MonoScale: Scaling Multi-Agent System with Monotonic Improvement
by: Shao, Shuai, et al.
Published: (2026)
by: Shao, Shuai, et al.
Published: (2026)
SCUBA: Salesforce Computer Use Benchmark
by: Dai, Yutong, et al.
Published: (2025)
by: Dai, Yutong, et al.
Published: (2025)
RvB: Automating AI System Hardening via Iterative Red-Blue Games
by: Huang, Lige, et al.
Published: (2026)
by: Huang, Lige, et al.
Published: (2026)
The Why Behind the Action: Unveiling Internal Drivers via Agentic Attribution
by: Qian, Chen, et al.
Published: (2026)
by: Qian, Chen, et al.
Published: (2026)
On the Reliability of Computer Use Agents
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2026)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2026)
Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents
by: Zhang, Boxuan, et al.
Published: (2025)
by: Zhang, Boxuan, et al.
Published: (2025)
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
by: Hu, Xuhao, et al.
Published: (2025)
by: Hu, Xuhao, et al.
Published: (2025)
Rethinking Entropy Regularization in Large Reasoning Models
by: Jiang, Yuxian, et al.
Published: (2025)
by: Jiang, Yuxian, et al.
Published: (2025)
Similar Items
-
OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents
by: Abhyankar, Reyna, et al.
Published: (2025) -
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
by: Xie, Tianbao, et al.
Published: (2024) -
macOSWorld: A Multilingual Interactive Benchmark for GUI Agents
by: Yang, Pei, et al.
Published: (2025) -
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
by: Shao, Shuai, et al.
Published: (2025) -
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
by: Li, Yu, et al.
Published: (2026)