Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bao, Muyi, Cai, Yuxin, Xu, Hang, Li, Zongtai, He, Jinxi, Tang, Jingfan, Lv, Chen, Zhang, Ji, Xie, Yaqi, Wang, Wenshan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UniGoal: Towards Universal Zero-shot Goal-oriented Navigation
von: Yin, Hang, et al.
Veröffentlicht: (2025)
von: Yin, Hang, et al.
Veröffentlicht: (2025)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
von: Liang, Wenqi, et al.
Veröffentlicht: (2025)
von: Liang, Wenqi, et al.
Veröffentlicht: (2025)
WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation
von: Nie, Dujun, et al.
Veröffentlicht: (2025)
von: Nie, Dujun, et al.
Veröffentlicht: (2025)
RSRNav: Reasoning Spatial Relationship for Image-Goal Navigation
von: Qin, Zheng, et al.
Veröffentlicht: (2025)
von: Qin, Zheng, et al.
Veröffentlicht: (2025)
VLD: Visual Language Goal Distance for Reinforcement Learning Navigation
von: Milikic, Lazar, et al.
Veröffentlicht: (2025)
von: Milikic, Lazar, et al.
Veröffentlicht: (2025)
Language-Based Augmentation to Address Shortcut Learning in Object Goal Navigation
von: Hoftijzer, Dennis, et al.
Veröffentlicht: (2024)
von: Hoftijzer, Dennis, et al.
Veröffentlicht: (2024)
LiteVLoc: Map-Lite Visual Localization for Image Goal Navigation
von: Jiao, Jianhao, et al.
Veröffentlicht: (2024)
von: Jiao, Jianhao, et al.
Veröffentlicht: (2024)
MOPA: Modular Object Navigation with PointGoal Agents
von: Raychaudhuri, Sonia, et al.
Veröffentlicht: (2023)
von: Raychaudhuri, Sonia, et al.
Veröffentlicht: (2023)
PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models
von: Wan, Jiansong, et al.
Veröffentlicht: (2025)
von: Wan, Jiansong, et al.
Veröffentlicht: (2025)
FOM-Nav: Frontier-Object Maps for Object Goal Navigation
von: Chabal, Thomas, et al.
Veröffentlicht: (2025)
von: Chabal, Thomas, et al.
Veröffentlicht: (2025)
TartanGround: A Large-Scale Dataset for Ground Robot Perception and Navigation
von: Patel, Manthan, et al.
Veröffentlicht: (2025)
von: Patel, Manthan, et al.
Veröffentlicht: (2025)
IRef-VLA: A Benchmark for Interactive Referential Grounding with Imperfect Language in 3D Scenes
von: Zhang, Haochen, et al.
Veröffentlicht: (2025)
von: Zhang, Haochen, et al.
Veröffentlicht: (2025)
Instance-aware Exploration-Verification-Exploitation for Instance ImageGoal Navigation
von: Lei, Xiaohan, et al.
Veröffentlicht: (2024)
von: Lei, Xiaohan, et al.
Veröffentlicht: (2024)
Cognitive Planning for Object Goal Navigation using Generative AI Models
von: S, Arjun P, et al.
Veröffentlicht: (2024)
von: S, Arjun P, et al.
Veröffentlicht: (2024)
LLM-RG: Referential Grounding in Outdoor Scenarios using Large Language Models
von: Saxena, Pranav, et al.
Veröffentlicht: (2025)
von: Saxena, Pranav, et al.
Veröffentlicht: (2025)
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
von: Lin, Yihan, et al.
Veröffentlicht: (2026)
von: Lin, Yihan, et al.
Veröffentlicht: (2026)
DiffPixelFormer: Differential Pixel-Aware Transformer for RGB-D Indoor Scene Segmentation
von: Gong, Yan, et al.
Veröffentlicht: (2025)
von: Gong, Yan, et al.
Veröffentlicht: (2025)
DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation
von: Fekri, Pedram, et al.
Veröffentlicht: (2025)
von: Fekri, Pedram, et al.
Veröffentlicht: (2025)
UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents
von: Xiao, Jianqiang, et al.
Veröffentlicht: (2025)
von: Xiao, Jianqiang, et al.
Veröffentlicht: (2025)
Hierarchical Scoring with 3D Gaussian Splatting for Instance Image-Goal Navigation
von: Deng, Yijie, et al.
Veröffentlicht: (2025)
von: Deng, Yijie, et al.
Veröffentlicht: (2025)
PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding
von: Wen, Junjie, et al.
Veröffentlicht: (2026)
von: Wen, Junjie, et al.
Veröffentlicht: (2026)
Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals
von: Gillman, Nate, et al.
Veröffentlicht: (2026)
von: Gillman, Nate, et al.
Veröffentlicht: (2026)
Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal Navigation
von: Li, Badi, et al.
Veröffentlicht: (2025)
von: Li, Badi, et al.
Veröffentlicht: (2025)
LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation
von: Miao, Bo, et al.
Veröffentlicht: (2026)
von: Miao, Bo, et al.
Veröffentlicht: (2026)
APEX: A Decoupled Memory-based Explorer for Asynchronous Aerial Object Goal Navigation
von: Zhang, Daoxuan, et al.
Veröffentlicht: (2026)
von: Zhang, Daoxuan, et al.
Veröffentlicht: (2026)
REST: Receding Horizon Explorative Steiner Tree for Zero-Shot Object-Goal Navigation
von: Xiao, Shuqi, et al.
Veröffentlicht: (2026)
von: Xiao, Shuqi, et al.
Veröffentlicht: (2026)
TDANet: Target-Directed Attention Network For Object-Goal Visual Navigation With Zero-Shot Ability
von: Lian, Shiwei, et al.
Veröffentlicht: (2024)
von: Lian, Shiwei, et al.
Veröffentlicht: (2024)
AnyImageNav: Any-View Geometry for Precise Last-Meter Image-Goal Navigation
von: Deng, Yijie, et al.
Veröffentlicht: (2026)
von: Deng, Yijie, et al.
Veröffentlicht: (2026)
Collision-Aware Object-Goal Visual Navigation via Two-Stage Deep Reinforcement Learning
von: Wang, Hongwu, et al.
Veröffentlicht: (2025)
von: Wang, Hongwu, et al.
Veröffentlicht: (2025)
CL-CoTNav: Closed-Loop Hierarchical Chain-of-Thought for Zero-Shot Object-Goal Navigation with Vision-Language Models
von: Cai, Yuxin, et al.
Veröffentlicht: (2025)
von: Cai, Yuxin, et al.
Veröffentlicht: (2025)
Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation
von: Chen, Bolei, et al.
Veröffentlicht: (2025)
von: Chen, Bolei, et al.
Veröffentlicht: (2025)
MPVO: Motion-Prior based Visual Odometry for PointGoal Navigation
von: Paul, Sayan, et al.
Veröffentlicht: (2024)
von: Paul, Sayan, et al.
Veröffentlicht: (2024)
Pixel Motion as Universal Representation for Robot Control
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2025)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2025)
EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding
von: Su, Yuejiao, et al.
Veröffentlicht: (2026)
von: Su, Yuejiao, et al.
Veröffentlicht: (2026)
GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation
von: Yin, Hang, et al.
Veröffentlicht: (2025)
von: Yin, Hang, et al.
Veröffentlicht: (2025)
W-PoseNet: Dense Correspondence Regularized Pixel Pair Pose Regression
von: Xu, Zelin, et al.
Veröffentlicht: (2019)
von: Xu, Zelin, et al.
Veröffentlicht: (2019)
CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs
von: Cao, Yihan, et al.
Veröffentlicht: (2024)
von: Cao, Yihan, et al.
Veröffentlicht: (2024)
AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer
von: Xue, Haoru, et al.
Veröffentlicht: (2025)
von: Xue, Haoru, et al.
Veröffentlicht: (2025)
Image-Goal Navigation Using Refined Feature Guidance and Scene Graph Enhancement
von: Feng, Zhicheng, et al.
Veröffentlicht: (2025)
von: Feng, Zhicheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
UniGoal: Towards Universal Zero-shot Goal-oriented Navigation
von: Yin, Hang, et al.
Veröffentlicht: (2025) -
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
von: Liang, Wenqi, et al.
Veröffentlicht: (2025) -
WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation
von: Nie, Dujun, et al.
Veröffentlicht: (2025) -
RSRNav: Reasoning Spatial Relationship for Image-Goal Navigation
von: Qin, Zheng, et al.
Veröffentlicht: (2025) -
VLD: Visual Language Goal Distance for Reinforcement Learning Navigation
von: Milikic, Lazar, et al.
Veröffentlicht: (2025)