AstraNav-Memory: Contexts Compression for Long Memory

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ren, Botao, Hu, Junjun, Xue, Xinda, Luo, Minghua, Chen, Jintao, Bai, Haochen, You, Liangliang, Xu, Mu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912789198536704
author Ren, Botao
Hu, Junjun
Xue, Xinda
Luo, Minghua
Chen, Jintao
Bai, Haochen
You, Liangliang
Xu, Mu
author_facet Ren, Botao
Hu, Junjun
Xue, Xinda
Luo, Minghua
Chen, Jintao
Bai, Haochen
You, Liangliang
Xu, Mu
contents Lifelong embodied navigation requires agents to accumulate, retain, and exploit spatial-semantic experience across tasks, enabling efficient exploration in novel environments and rapid goal reaching in familiar ones. While object-centric memory is interpretable, it depends on detection and reconstruction pipelines that limit robustness and scalability. We propose an image-centric memory framework that achieves long-term implicit memory via an efficient visual context compression module end-to-end coupled with a Qwen2.5-VL-based navigation policy. Built atop a ViT backbone with frozen DINOv3 features and lightweight PixelUnshuffle+Conv blocks, our visual tokenizer supports configurable compression rates; for example, under a representative 16$\times$ compression setting, each image is encoded with about 30 tokens, expanding the effective context capacity from tens to hundreds of images. Experimental results on GOAT-Bench and HM3D-OVON show that our method achieves state-of-the-art navigation performance, improving exploration in unfamiliar environments and shortening paths in familiar ones. Ablation studies further reveal that moderate compression provides the best balance between efficiency and accuracy. These findings position compressed image-centric memory as a practical and scalable interface for lifelong embodied agents, enabling them to reason over long visual histories and navigate with human-like efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21627
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AstraNav-Memory: Contexts Compression for Long Memory
Ren, Botao
Hu, Junjun
Xue, Xinda
Luo, Minghua
Chen, Jintao
Bai, Haochen
You, Liangliang
Xu, Mu
Robotics
Lifelong embodied navigation requires agents to accumulate, retain, and exploit spatial-semantic experience across tasks, enabling efficient exploration in novel environments and rapid goal reaching in familiar ones. While object-centric memory is interpretable, it depends on detection and reconstruction pipelines that limit robustness and scalability. We propose an image-centric memory framework that achieves long-term implicit memory via an efficient visual context compression module end-to-end coupled with a Qwen2.5-VL-based navigation policy. Built atop a ViT backbone with frozen DINOv3 features and lightweight PixelUnshuffle+Conv blocks, our visual tokenizer supports configurable compression rates; for example, under a representative 16$\times$ compression setting, each image is encoded with about 30 tokens, expanding the effective context capacity from tens to hundreds of images. Experimental results on GOAT-Bench and HM3D-OVON show that our method achieves state-of-the-art navigation performance, improving exploration in unfamiliar environments and shortening paths in familiar ones. Ablation studies further reveal that moderate compression provides the best balance between efficiency and accuracy. These findings position compressed image-centric memory as a practical and scalable interface for lifelong embodied agents, enabling them to reason over long visual histories and navigate with human-like efficiency.
title AstraNav-Memory: Contexts Compression for Long Memory
topic Robotics
url https://arxiv.org/abs/2512.21627