Implicit Geometry Representations for Vision-and-Language Navigation from Web Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Mingfei, Hao, Haihong, Ma, Liang, Zhumakhanova, Kamila, Radionova, Ekaterina, Zhang, Jingyi, Chang, Xiaojun, Liang, Xiaodan, Laptev, Ivan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation
von: Han, Mingfei, et al.
Veröffentlicht: (2024)
von: Han, Mingfei, et al.
Veröffentlicht: (2024)
GLaD: Geometric Latent Distillation for Vision-Language-Action Models
von: Guo, Minghao, et al.
Veröffentlicht: (2025)
von: Guo, Minghao, et al.
Veröffentlicht: (2025)
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
von: Hao, Haihong, et al.
Veröffentlicht: (2026)
von: Hao, Haihong, et al.
Veröffentlicht: (2026)
NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning
von: Lin, Bingqian, et al.
Veröffentlicht: (2024)
von: Lin, Bingqian, et al.
Veröffentlicht: (2024)
Choose What to Observe: Task-Aware Semantic-Geometric Representations for Visuomotor Policy
von: Ding, Haoran, et al.
Veröffentlicht: (2026)
von: Ding, Haoran, et al.
Veröffentlicht: (2026)
MALMM: Multi-Agent Large Language Models for Zero-Shot Robotics Manipulation
von: Singh, Harsh, et al.
Veröffentlicht: (2024)
von: Singh, Harsh, et al.
Veröffentlicht: (2024)
Affordances-Oriented Planning using Foundation Models for Continuous Vision-Language Navigation
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models
von: Yan, Yu, et al.
Veröffentlicht: (2024)
von: Yan, Yu, et al.
Veröffentlicht: (2024)
JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation
von: Zeng, Shuang, et al.
Veröffentlicht: (2025)
von: Zeng, Shuang, et al.
Veröffentlicht: (2025)
Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) \rightarrow G$): Vision-Geometry Backbones over Language and Video Models
von: Song, Zijian, et al.
Veröffentlicht: (2026)
von: Song, Zijian, et al.
Veröffentlicht: (2026)
PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly
von: Ma, Liang, et al.
Veröffentlicht: (2025)
von: Ma, Liang, et al.
Veröffentlicht: (2025)
CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
von: Hao, Haihong, et al.
Veröffentlicht: (2025)
von: Hao, Haihong, et al.
Veröffentlicht: (2025)
VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation
von: Wen, Youpeng, et al.
Veröffentlicht: (2024)
von: Wen, Youpeng, et al.
Veröffentlicht: (2024)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
von: Liang, Wenqi, et al.
Veröffentlicht: (2025)
von: Liang, Wenqi, et al.
Veröffentlicht: (2025)
GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
von: Abouzeid, Ali, et al.
Veröffentlicht: (2025)
von: Abouzeid, Ali, et al.
Veröffentlicht: (2025)
Unseen from Seen: Rewriting Observation-Instruction Using Foundation Models for Augmenting Vision-Language Navigation
von: Wei, Ziming, et al.
Veröffentlicht: (2025)
von: Wei, Ziming, et al.
Veröffentlicht: (2025)
A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model
von: Zhang, Kaidong, et al.
Veröffentlicht: (2026)
von: Zhang, Kaidong, et al.
Veröffentlicht: (2026)
MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation
von: Dai, Tingjun, et al.
Veröffentlicht: (2026)
von: Dai, Tingjun, et al.
Veröffentlicht: (2026)
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
von: Chen, Zerui, et al.
Veröffentlicht: (2024)
von: Chen, Zerui, et al.
Veröffentlicht: (2024)
Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation
von: Jin, Minghao, et al.
Veröffentlicht: (2026)
von: Jin, Minghao, et al.
Veröffentlicht: (2026)
AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation
von: Ding, Xin, et al.
Veröffentlicht: (2025)
von: Ding, Xin, et al.
Veröffentlicht: (2025)
LAMP: Implicit Language Map for Robot Navigation
von: Lee, Sibaek, et al.
Veröffentlicht: (2026)
von: Lee, Sibaek, et al.
Veröffentlicht: (2026)
Exploring Spatial Representation to Enhance LLM Reasoning in Aerial Vision-Language Navigation
von: Gao, Yunpeng, et al.
Veröffentlicht: (2024)
von: Gao, Yunpeng, et al.
Veröffentlicht: (2024)
DyGeoVLN: Infusing Dynamic Geometry Foundation Model into Vision-Language Navigation
von: Liu, Xiangchen, et al.
Veröffentlicht: (2026)
von: Liu, Xiangchen, et al.
Veröffentlicht: (2026)
Sparse Video Generation Propels Real-World Beyond-the-View Vision-Language Navigation
von: Zhang, Hai, et al.
Veröffentlicht: (2026)
von: Zhang, Hai, et al.
Veröffentlicht: (2026)
Off-Road Navigation via Implicit Neural Representation of Terrain Traversability
von: Jia, Yixuan, et al.
Veröffentlicht: (2025)
von: Jia, Yixuan, et al.
Veröffentlicht: (2025)
History-Conditioned Spatio-Temporal Visual Token Pruning for Efficient Vision-Language Navigation
von: Wang, Qitong, et al.
Veröffentlicht: (2026)
von: Wang, Qitong, et al.
Veröffentlicht: (2026)
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
von: Li, Fuhao, et al.
Veröffentlicht: (2025)
von: Li, Fuhao, et al.
Veröffentlicht: (2025)
MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2025)
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2025)
Enhancing Exploratory Capability of Visual Navigation Using Uncertainty of Implicit Scene Representation
von: Wang, Yichen, et al.
Veröffentlicht: (2024)
von: Wang, Yichen, et al.
Veröffentlicht: (2024)
Safe-VLN: Collision Avoidance for Vision-and-Language Navigation of Autonomous Robots Operating in Continuous Environments
von: Yue, Lu, et al.
Veröffentlicht: (2023)
von: Yue, Lu, et al.
Veröffentlicht: (2023)
AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment
von: Kong, Weijie, et al.
Veröffentlicht: (2026)
von: Kong, Weijie, et al.
Veröffentlicht: (2026)
BLAZER: Bootstrapping LLM-based Manipulation Agents with Zero-Shot Data Generation
von: Das, Rocktim Jyoti, et al.
Veröffentlicht: (2025)
von: Das, Rocktim Jyoti, et al.
Veröffentlicht: (2025)
Cross-Hand Latent Representation for Vision-Language-Action Models
von: Jiang, Guangqi, et al.
Veröffentlicht: (2026)
von: Jiang, Guangqi, et al.
Veröffentlicht: (2026)
FloorPlan-VLN: A New Paradigm for Floor Plan Guided Vision-Language Navigation
von: Chen, Kehan, et al.
Veröffentlicht: (2026)
von: Chen, Kehan, et al.
Veröffentlicht: (2026)
Learning feasible transitions for efficient contact planning
von: Akizhanov, Rikhat, et al.
Veröffentlicht: (2024)
von: Akizhanov, Rikhat, et al.
Veröffentlicht: (2024)
Structured Observation Language for Efficient and Generalizable Vision-Language Navigation
von: Peng, Daojie, et al.
Veröffentlicht: (2026)
von: Peng, Daojie, et al.
Veröffentlicht: (2026)
VLM-Social-Nav: Socially Aware Robot Navigation through Scoring using Vision-Language Models
von: Song, Daeun, et al.
Veröffentlicht: (2024)
von: Song, Daeun, et al.
Veröffentlicht: (2024)
A Deployable Embodied Vision-Language Navigation System with Hierarchical Cognition and Context-Aware Exploration
von: Xu, Kuan, et al.
Veröffentlicht: (2026)
von: Xu, Kuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation
von: Han, Mingfei, et al.
Veröffentlicht: (2024) -
GLaD: Geometric Latent Distillation for Vision-Language-Action Models
von: Guo, Minghao, et al.
Veröffentlicht: (2025) -
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
von: Hao, Haihong, et al.
Veröffentlicht: (2026) -
NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning
von: Lin, Bingqian, et al.
Veröffentlicht: (2024) -
Choose What to Observe: Task-Aware Semantic-Geometric Representations for Visuomotor Policy
von: Ding, Haoran, et al.
Veröffentlicht: (2026)