Saved in:
Bibliographic Details
Main Authors: Luo, Zhihao, Yan, Wentao, Gong, Jingyu, Wang, Min, Zhang, Zhizhong, Wang, Xuhong, Xie, Yuan, Tan, Xin
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.02046
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918477634207744
author Luo, Zhihao
Yan, Wentao
Gong, Jingyu
Wang, Min
Zhang, Zhizhong
Wang, Xuhong
Xie, Yuan
Tan, Xin
author_facet Luo, Zhihao
Yan, Wentao
Gong, Jingyu
Wang, Min
Zhang, Zhizhong
Wang, Xuhong
Xie, Yuan
Tan, Xin
contents Recent advances in Graphical User Interface (GUI) and embodied navigation have driven progress, yet these domains have largely evolved in isolation, with disparate datasets and training paradigms. In this paper, we observe that both tasks can be formulated as Markov Decision Processes (MDP), suggesting a foundational principle for their unification. Hence, we present NaviMaster, the first unified agent capable of unifying GUI navigation and embodied navigation within a single framework. Specifically, NaviMaster (i) proposes a visual-target trajectory collection pipeline that generates trajectories for both GUI and embodied tasks using a single formulation. (ii) employs a unified reinforcement learning framework on the mix data to improve generalization. (iii) designs a novel distance-aware reward to ensure efficient learning from the trajectories. Through extensive experiments on out-of-domain benchmarks, NaviMaster is shown to outperform state-of-the-art agents in GUI navigation, spatial affordance prediction, and embodied navigation. Ablation studies further demonstrate the efficacy of our unified training strategy, data mixing strategy, and reward design. Our codes, data, and checkpoints are available at https://iron-boyy.github.io/navimaster-page/.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02046
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks
Luo, Zhihao
Yan, Wentao
Gong, Jingyu
Wang, Min
Zhang, Zhizhong
Wang, Xuhong
Xie, Yuan
Tan, Xin
Robotics
Machine Learning
Recent advances in Graphical User Interface (GUI) and embodied navigation have driven progress, yet these domains have largely evolved in isolation, with disparate datasets and training paradigms. In this paper, we observe that both tasks can be formulated as Markov Decision Processes (MDP), suggesting a foundational principle for their unification. Hence, we present NaviMaster, the first unified agent capable of unifying GUI navigation and embodied navigation within a single framework. Specifically, NaviMaster (i) proposes a visual-target trajectory collection pipeline that generates trajectories for both GUI and embodied tasks using a single formulation. (ii) employs a unified reinforcement learning framework on the mix data to improve generalization. (iii) designs a novel distance-aware reward to ensure efficient learning from the trajectories. Through extensive experiments on out-of-domain benchmarks, NaviMaster is shown to outperform state-of-the-art agents in GUI navigation, spatial affordance prediction, and embodied navigation. Ablation studies further demonstrate the efficacy of our unified training strategy, data mixing strategy, and reward design. Our codes, data, and checkpoints are available at https://iron-boyy.github.io/navimaster-page/.
title NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks
topic Robotics
Machine Learning
url https://arxiv.org/abs/2508.02046