AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Jeonghyeon, Joung, Byeongjun, Lee, Junwon, Lee, Joohyung, Min, Taehoon, Lee, Sunjae
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915950318583808
author Kim, Jeonghyeon
Joung, Byeongjun
Lee, Junwon
Lee, Joohyung
Min, Taehoon
Lee, Sunjae
author_facet Kim, Jeonghyeon
Joung, Byeongjun
Lee, Junwon
Lee, Joohyung
Min, Taehoon
Lee, Sunjae
contents Mobile GUI agents can automate smartphone tasks by interacting directly with app interfaces, but how they should communicate with users during execution remains underexplored. Existing systems rely on two extremes: foreground execution, which maximizes transparency but prevents multitasking, and background execution, which supports multitasking but provides little visual awareness. Through iterative formative studies, we found that users prefer a hybrid model with just-in-time visual interaction, but the most effective visualization modality depends on the task. Motivated by this, we present AgentLens, a mobile GUI agent that adaptively uses three visual modalities during human-agent interaction: Full UI, Partial UI, and GenUI. AgentLens extends a standard mobile agent with adaptive communication actions and uses Virtual Display to enable background execution with selective visual overlays. In a controlled study with 21 participants, AgentLens was preferred by 85.7% of participants and achieved the highest usability (1.94 Overall PSSUQ) and adoption-intent (6.43/7).
format Preprint
id arxiv_https___arxiv_org_abs_2604_20279
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents
Kim, Jeonghyeon
Joung, Byeongjun
Lee, Junwon
Lee, Joohyung
Min, Taehoon
Lee, Sunjae
Human-Computer Interaction
Artificial Intelligence
Multiagent Systems
Mobile GUI agents can automate smartphone tasks by interacting directly with app interfaces, but how they should communicate with users during execution remains underexplored. Existing systems rely on two extremes: foreground execution, which maximizes transparency but prevents multitasking, and background execution, which supports multitasking but provides little visual awareness. Through iterative formative studies, we found that users prefer a hybrid model with just-in-time visual interaction, but the most effective visualization modality depends on the task. Motivated by this, we present AgentLens, a mobile GUI agent that adaptively uses three visual modalities during human-agent interaction: Full UI, Partial UI, and GenUI. AgentLens extends a standard mobile agent with adaptive communication actions and uses Virtual Display to enable background execution with selective visual overlays. In a controlled study with 21 participants, AgentLens was preferred by 85.7% of participants and achieved the highest usability (1.94 Overall PSSUQ) and adoption-intent (6.43/7).
title AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents
topic Human-Computer Interaction
Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2604.20279