Image-Goal Navigation Using Refined Feature Guidance and Scene Graph Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Zhicheng, Chen, Xieyuanli, Shi, Chenghao, Luo, Lun, Chen, Zhichao, Liu, Yun-Hui, Lu, Huimin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910875590328320
author Feng, Zhicheng
Chen, Xieyuanli
Shi, Chenghao
Luo, Lun
Chen, Zhichao
Liu, Yun-Hui
Lu, Huimin
author_facet Feng, Zhicheng
Chen, Xieyuanli
Shi, Chenghao
Luo, Lun
Chen, Zhichao
Liu, Yun-Hui
Lu, Huimin
contents In this paper, we introduce a novel image-goal navigation approach, named RFSG. Our focus lies in leveraging the fine-grained connections between goals, observations, and the environment within limited image data, all the while keeping the navigation architecture simple and lightweight. To this end, we propose the spatial-channel attention mechanism, enabling the network to learn the importance of multi-dimensional features to fuse the goal and observation features. In addition, a selfdistillation mechanism is incorporated to further enhance the feature representation capabilities. Given that the navigation task needs surrounding environmental information for more efficient navigation, we propose an image scene graph to establish feature associations at both the image and object levels, effectively encoding the surrounding scene information. Crossscene performance validation was conducted on the Gibson and HM3D datasets, and the proposed method achieved stateof-the-art results among mainstream methods, with a speed of up to 53.5 frames per second on an RTX3080. This contributes to the realization of end-to-end image-goal navigation in realworld scenarios. The implementation and model of our method have been released at: https://github.com/nubot-nudt/RFSG.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10986
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Image-Goal Navigation Using Refined Feature Guidance and Scene Graph Enhancement
Feng, Zhicheng
Chen, Xieyuanli
Shi, Chenghao
Luo, Lun
Chen, Zhichao
Liu, Yun-Hui
Lu, Huimin
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
In this paper, we introduce a novel image-goal navigation approach, named RFSG. Our focus lies in leveraging the fine-grained connections between goals, observations, and the environment within limited image data, all the while keeping the navigation architecture simple and lightweight. To this end, we propose the spatial-channel attention mechanism, enabling the network to learn the importance of multi-dimensional features to fuse the goal and observation features. In addition, a selfdistillation mechanism is incorporated to further enhance the feature representation capabilities. Given that the navigation task needs surrounding environmental information for more efficient navigation, we propose an image scene graph to establish feature associations at both the image and object levels, effectively encoding the surrounding scene information. Crossscene performance validation was conducted on the Gibson and HM3D datasets, and the proposed method achieved stateof-the-art results among mainstream methods, with a speed of up to 53.5 frames per second on an RTX3080. This contributes to the realization of end-to-end image-goal navigation in realworld scenarios. The implementation and model of our method have been released at: https://github.com/nubot-nudt/RFSG.
title Image-Goal Navigation Using Refined Feature Guidance and Scene Graph Enhancement
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.10986