Improving Random Testing via LLM-powered UI Tarpit Escaping for Mobile Apps

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Mengqian, Xiong, Yiheng, Chang, Le, Su, Ting, Wan, Chengcheng, Miao, Weikai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911586534293504
author Xu, Mengqian
Xiong, Yiheng
Chang, Le
Su, Ting
Wan, Chengcheng
Miao, Weikai
author_facet Xu, Mengqian
Xiong, Yiheng
Chang, Le
Su, Ting
Wan, Chengcheng
Miao, Weikai
contents Random GUI testing is a widely-used technique for testing mobile apps. However, its effectiveness is limited by the notorious issue -- UI exploration tarpits, where the exploration is trapped in local UI regions, thus impeding test coverage and bug discovery. In this experience paper, we introduce LLM-powered random GUI Testing, a novel hybrid testing approach to mitigating UI tarpits during random testing. Our approach monitors UI similarity to identify tarpits and query LLMs to suggest promising events for escaping the encountered tarpits. We implement our approach on top of two different automated input generation (AIG) tools for mobile apps: (1) HybridMonkey upon Monkey, a state-of-the-practice tool; and (2) HybridDroidbot upon Droidbot, a state-of-the-art tool. We evaluated them on 12 popular, real-world apps. The results show that HybridMonkey and HybridDroidbot outperform all baselines, achieving average coverage improvements of 54.8% and 44.8%, respectively, and detecting the highest number of unique crashes. In total, we found 75 unique bugs, including 34 previously unknown bugs. To date, 26 bugs have been confirmed and fixed. We also applied HybridMonkey on WeChat, a popular industrial app with billions of monthly active users. HybridMonkey achieved higher activity coverage and found more bugs than random testing.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06763
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Improving Random Testing via LLM-powered UI Tarpit Escaping for Mobile Apps
Xu, Mengqian
Xiong, Yiheng
Chang, Le
Su, Ting
Wan, Chengcheng
Miao, Weikai
Software Engineering
Random GUI testing is a widely-used technique for testing mobile apps. However, its effectiveness is limited by the notorious issue -- UI exploration tarpits, where the exploration is trapped in local UI regions, thus impeding test coverage and bug discovery. In this experience paper, we introduce LLM-powered random GUI Testing, a novel hybrid testing approach to mitigating UI tarpits during random testing. Our approach monitors UI similarity to identify tarpits and query LLMs to suggest promising events for escaping the encountered tarpits. We implement our approach on top of two different automated input generation (AIG) tools for mobile apps: (1) HybridMonkey upon Monkey, a state-of-the-practice tool; and (2) HybridDroidbot upon Droidbot, a state-of-the-art tool. We evaluated them on 12 popular, real-world apps. The results show that HybridMonkey and HybridDroidbot outperform all baselines, achieving average coverage improvements of 54.8% and 44.8%, respectively, and detecting the highest number of unique crashes. In total, we found 75 unique bugs, including 34 previously unknown bugs. To date, 26 bugs have been confirmed and fixed. We also applied HybridMonkey on WeChat, a popular industrial app with billions of monthly active users. HybridMonkey achieved higher activity coverage and found more bugs than random testing.
title Improving Random Testing via LLM-powered UI Tarpit Escaping for Mobile Apps
topic Software Engineering
url https://arxiv.org/abs/2604.06763