When Engineering Outruns Intelligence: Rethinking Instruction-Guided Navigation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Aghaei, Matin, Zhang, Lingfeng, Alomrani, Mohammad Ali, Biparva, Mahdi, Zhang, Yingxue
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918485489090560
author Aghaei, Matin
Zhang, Lingfeng
Alomrani, Mohammad Ali
Biparva, Mahdi
Zhang, Yingxue
author_facet Aghaei, Matin
Zhang, Lingfeng
Alomrani, Mohammad Ali
Biparva, Mahdi
Zhang, Yingxue
contents Recent ObjectNav systems credit large language models (LLMs) for sizable zero-shot gains, yet it remains unclear how much comes from language versus geometry. We revisit this question by re-evaluating an instruction-guided pipeline, InstructNav, under a detector-controlled setting and introducing two training-free variants that only alter the action value map: a geometry-only Frontier Proximity Explorer (FPE) and a lightweight Semantic-Heuristic Frontier (SHF) that polls the LLM with simple frontier votes. Across HM3D and MP3D, FPE matches or exceeds the detector-controlled instruction follower while using no API calls and running faster; SHF attains comparable accuracy with a smaller, localized language prior. These results suggest that carefully engineered frontier geometry accounts for much of the reported progress, and that language is most reliable as a light heuristic rather than an end-to-end planner. Code available at: https://github.com/matinaghaei/instructnav-scrutinized
format Preprint
id arxiv_https___arxiv_org_abs_2507_20021
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Engineering Outruns Intelligence: Rethinking Instruction-Guided Navigation
Aghaei, Matin
Zhang, Lingfeng
Alomrani, Mohammad Ali
Biparva, Mahdi
Zhang, Yingxue
Robotics
Artificial Intelligence
Machine Learning
Recent ObjectNav systems credit large language models (LLMs) for sizable zero-shot gains, yet it remains unclear how much comes from language versus geometry. We revisit this question by re-evaluating an instruction-guided pipeline, InstructNav, under a detector-controlled setting and introducing two training-free variants that only alter the action value map: a geometry-only Frontier Proximity Explorer (FPE) and a lightweight Semantic-Heuristic Frontier (SHF) that polls the LLM with simple frontier votes. Across HM3D and MP3D, FPE matches or exceeds the detector-controlled instruction follower while using no API calls and running faster; SHF attains comparable accuracy with a smaller, localized language prior. These results suggest that carefully engineered frontier geometry accounts for much of the reported progress, and that language is most reliable as a light heuristic rather than an end-to-end planner. Code available at: https://github.com/matinaghaei/instructnav-scrutinized
title When Engineering Outruns Intelligence: Rethinking Instruction-Guided Navigation
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.20021