Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Zhengyang, Zhang, Yi, Li, Chenxin, Lai, Xin, Lyu, Pengyuan, Guo, Yiduo, Wang, Weinong, Li, Junyi, Ding, Yang, Shen, Huawen, Fang, Zhengyao, Zhou, Xingran, Wu, Liang, Tang, Fei, Fan, Sunqi, Peng, Shangpin, Ruan, Zheng, Zhang, Anran, Wang, Benyou, Zhang, Chengquan, Hu, Han
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911662151303168
author Tang, Zhengyang
Zhang, Yi
Li, Chenxin
Lai, Xin
Lyu, Pengyuan
Guo, Yiduo
Wang, Weinong
Li, Junyi
Ding, Yang
Shen, Huawen
Fang, Zhengyao
Zhou, Xingran
Wu, Liang
Tang, Fei
Fan, Sunqi
Peng, Shangpin
Ruan, Zheng
Zhang, Anran
Wang, Benyou
Zhang, Chengquan
Hu, Han
author_facet Tang, Zhengyang
Zhang, Yi
Li, Chenxin
Lai, Xin
Lyu, Pengyuan
Guo, Yiduo
Wang, Weinong
Li, Junyi
Ding, Yang
Shen, Huawen
Fang, Zhengyao
Zhou, Xingran
Wu, Liang
Tang, Fei
Fan, Sunqi
Peng, Shangpin
Ruan, Zheng
Zhang, Anran
Wang, Benyou
Zhang, Chengquan
Hu, Han
contents When a phone-use agent avoids harm, does that show safety, or simply inability to act? Existing evaluations often cannot tell. A harmful outcome may be avoided because the agent recognized the risk and chose the safe action, or because it failed to understand the screen or execute any relevant action at all. These cases have different causes and call for different fixes, yet current benchmarks often merge them under task success, refusal, or final harmful outcome. We address this problem with PhoneSafety, a benchmark of 700 safety-critical moments drawn from real phone interactions across more than 130 apps. Each instance isolates the next decision at a risky moment and asks a simple question: does the model take the safe action, take the unsafe action, or fail to do anything useful? We evaluate eight representative phone-use agents under this framework. Our results reveal two main patterns. First, stronger general phone-use ability does not reliably imply safer choices at risky moments. Models that perform better on ordinary app tasks are not always the ones that behave more safely when the next action matters. Second, failures to do anything useful behave like a capability signal rather than a safety signal: they are concentrated in more visually and operationally demanding settings and remain stable when the evaluation protocol changes. Across models, failures split into two recurring patterns: unsafe choices in settings where the model can act but chooses wrongly, and inability to act in more visually and operationally demanding screens. Overall, a harmless outcome is not enough to count as evidence of safety. Evaluating phone-use agents requires separating unsafe judgment from inability to act.
format Preprint
id arxiv_https___arxiv_org_abs_2605_07630
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents
Tang, Zhengyang
Zhang, Yi
Li, Chenxin
Lai, Xin
Lyu, Pengyuan
Guo, Yiduo
Wang, Weinong
Li, Junyi
Ding, Yang
Shen, Huawen
Fang, Zhengyao
Zhou, Xingran
Wu, Liang
Tang, Fei
Fan, Sunqi
Peng, Shangpin
Ruan, Zheng
Zhang, Anran
Wang, Benyou
Zhang, Chengquan
Hu, Han
Computation and Language
Artificial Intelligence
Machine Learning
When a phone-use agent avoids harm, does that show safety, or simply inability to act? Existing evaluations often cannot tell. A harmful outcome may be avoided because the agent recognized the risk and chose the safe action, or because it failed to understand the screen or execute any relevant action at all. These cases have different causes and call for different fixes, yet current benchmarks often merge them under task success, refusal, or final harmful outcome. We address this problem with PhoneSafety, a benchmark of 700 safety-critical moments drawn from real phone interactions across more than 130 apps. Each instance isolates the next decision at a risky moment and asks a simple question: does the model take the safe action, take the unsafe action, or fail to do anything useful? We evaluate eight representative phone-use agents under this framework. Our results reveal two main patterns. First, stronger general phone-use ability does not reliably imply safer choices at risky moments. Models that perform better on ordinary app tasks are not always the ones that behave more safely when the next action matters. Second, failures to do anything useful behave like a capability signal rather than a safety signal: they are concentrated in more visually and operationally demanding settings and remain stable when the evaluation protocol changes. Across models, failures split into two recurring patterns: unsafe choices in settings where the model can act but chooses wrongly, and inability to act in more visually and operationally demanding screens. Overall, a harmless outcome is not enough to count as evidence of safety. Evaluating phone-use agents requires separating unsafe judgment from inability to act.
title Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.07630