AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Yibo, Li, Jungang, Zhang, Linghao, Dongfang, Zihao, Wu, Biao, Tao, Sicheng, Yan, Yibo, Qin, Chenxi, Liu, Weiting, Lin, Zhixin, Li, Hanqian, Huang, Yu, Dai, Song, Hei, Yonghua, Ding, Yue, Li, Xiang, Wang, Shikang, Xu, Chengdong, Liu, Jingqi, Ma, Xueying, Zheng, Zhiwen, Zhang, Xiaofei, Wang, Bincheng, Yang, Nichen, Wu, Jie, Tian, Lihua, Li, Chen, Hu, Xuming
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912973750009856
author Shi, Yibo
Li, Jungang
Zhang, Linghao
Dongfang, Zihao
Wu, Biao
Tao, Sicheng
Yan, Yibo
Qin, Chenxi
Liu, Weiting
Lin, Zhixin
Li, Hanqian
Huang, Yu
Dai, Song
Hei, Yonghua
Ding, Yue
Li, Xiang
Wang, Shikang
Xu, Chengdong
Liu, Jingqi
Ma, Xueying
Zheng, Zhiwen
Zhang, Xiaofei
Wang, Bincheng
Yang, Nichen
Wu, Jie
Tian, Lihua
Li, Chen
Hu, Xuming
author_facet Shi, Yibo
Li, Jungang
Zhang, Linghao
Dongfang, Zihao
Wu, Biao
Tao, Sicheng
Yan, Yibo
Qin, Chenxi
Liu, Weiting
Lin, Zhixin
Li, Hanqian
Huang, Yu
Dai, Song
Hei, Yonghua
Ding, Yue
Li, Xiang
Wang, Shikang
Xu, Chengdong
Liu, Jingqi
Ma, Xueying
Zheng, Zhiwen
Zhang, Xiaofei
Wang, Bincheng
Yang, Nichen
Wu, Jie
Tian, Lihua
Li, Chen
Hu, Xuming
contents Long-horizon GUI agents are a key step toward real-world deployment, yet effective interaction memory under prevailing paradigms remains under-explored. Replaying full interaction sequences is redundant and amplifies noise, while summaries often erase dependency-critical information and traceability. We present AndroTMem, a diagnostic framework for anchored memory in long-horizon Android GUI agents. Its core benchmark, AndroTMem-Bench, comprises 1,069 tasks with 34,473 interaction steps (avg. 32.1 per task, max. 65). We evaluate agents with TCR (Task Complete Rate), focusing on tasks whose completion requires carrying forward critical intermediate state; AndroTMem-Bench is designed to enforce strong step-to-step causal dependencies, making sparse yet essential intermediate states decisive for downstream actions and centering interaction memory in evaluation. Across open- and closed-source GUI agents, we observe a consistent pattern: as interaction sequences grow longer, performance drops are driven mainly by within-task memory failures, not isolated perception errors or local action mistakes. Guided by this diagnosis, we propose Anchored State Memory (ASM), which represents interaction sequences as a compact set of causally linked intermediate-state anchors to enable subgoal-targeted retrieval and attribution-aware decision making. Across multiple settings and 12 evaluated GUI agents, ASM consistently outperforms full-sequence replay and summary-based baselines, improving TCR by 5%-30.16% and AMS by 4.93%-24.66%, indicating that anchored, structured memory effectively mitigates the interaction-memory bottleneck in long-horizon GUI tasks. The code, benchmark, and related resources are publicly available at [https://github.com/CVC2233/AndroTMem](https://github.com/CVC2233/AndroTMem).
format Preprint
id arxiv_https___arxiv_org_abs_2603_18429
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI Agents
Shi, Yibo
Li, Jungang
Zhang, Linghao
Dongfang, Zihao
Wu, Biao
Tao, Sicheng
Yan, Yibo
Qin, Chenxi
Liu, Weiting
Lin, Zhixin
Li, Hanqian
Huang, Yu
Dai, Song
Hei, Yonghua
Ding, Yue
Li, Xiang
Wang, Shikang
Xu, Chengdong
Liu, Jingqi
Ma, Xueying
Zheng, Zhiwen
Zhang, Xiaofei
Wang, Bincheng
Yang, Nichen
Wu, Jie
Tian, Lihua
Li, Chen
Hu, Xuming
Computer Vision and Pattern Recognition
Long-horizon GUI agents are a key step toward real-world deployment, yet effective interaction memory under prevailing paradigms remains under-explored. Replaying full interaction sequences is redundant and amplifies noise, while summaries often erase dependency-critical information and traceability. We present AndroTMem, a diagnostic framework for anchored memory in long-horizon Android GUI agents. Its core benchmark, AndroTMem-Bench, comprises 1,069 tasks with 34,473 interaction steps (avg. 32.1 per task, max. 65). We evaluate agents with TCR (Task Complete Rate), focusing on tasks whose completion requires carrying forward critical intermediate state; AndroTMem-Bench is designed to enforce strong step-to-step causal dependencies, making sparse yet essential intermediate states decisive for downstream actions and centering interaction memory in evaluation. Across open- and closed-source GUI agents, we observe a consistent pattern: as interaction sequences grow longer, performance drops are driven mainly by within-task memory failures, not isolated perception errors or local action mistakes. Guided by this diagnosis, we propose Anchored State Memory (ASM), which represents interaction sequences as a compact set of causally linked intermediate-state anchors to enable subgoal-targeted retrieval and attribution-aware decision making. Across multiple settings and 12 evaluated GUI agents, ASM consistently outperforms full-sequence replay and summary-based baselines, improving TCR by 5%-30.16% and AMS by 4.93%-24.66%, indicating that anchored, structured memory effectively mitigates the interaction-memory bottleneck in long-horizon GUI tasks. The code, benchmark, and related resources are publicly available at [https://github.com/CVC2233/AndroTMem](https://github.com/CVC2233/AndroTMem).
title AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI Agents
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.18429