Embodied Navigation Foundation Model
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jiazhao, Li, Anqi, Qi, Yunpeng, Li, Minghan, Liu, Jiahang, Wang, Shaoan, Liu, Haoran, Zhou, Gengze, Wu, Yuze, Li, Xingxing, Fan, Yuxin, Li, Wenjun, Chen, Zhibo, Gao, Fei, Wu, Qi, Zhang, Zhizheng, Wang, He |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TrackVLA: Embodied Visual Tracking in the Wild
by: Wang, Shaoan, et al.
Published: (2025)
by: Wang, Shaoan, et al.
Published: (2025)
TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
by: Liu, Jiahang, et al.
Published: (2025)
by: Liu, Jiahang, et al.
Published: (2025)
NavGSim: High-Fidelity Gaussian Splatting Simulator for Large-Scale Navigation
by: Liu, Jiahang, et al.
Published: (2026)
by: Liu, Jiahang, et al.
Published: (2026)
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
by: Zhang, Jiazhao, et al.
Published: (2024)
by: Zhang, Jiazhao, et al.
Published: (2024)
UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
by: Li, Anqi, et al.
Published: (2025)
by: Li, Anqi, et al.
Published: (2025)
SPAN-Nav: Generalized Spatial Awareness for Versatile Vision-Language Navigation
by: Liu, Jiahang, et al.
Published: (2026)
by: Liu, Jiahang, et al.
Published: (2026)
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
by: Zhang, Jiazhao, et al.
Published: (2024)
by: Zhang, Jiazhao, et al.
Published: (2024)
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
by: Xu, Tianyu, et al.
Published: (2025)
by: Xu, Tianyu, et al.
Published: (2025)
VLA-AN: An Efficient and Onboard Vision-Language-Action Framework for Aerial Navigation in Complex Environments
by: Wu, Yuze, et al.
Published: (2025)
by: Wu, Yuze, et al.
Published: (2025)
VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation agents
by: Zhao, Xunyi, et al.
Published: (2025)
by: Zhao, Xunyi, et al.
Published: (2025)
OctoNav: Towards Generalist Embodied Navigation
by: Gao, Chen, et al.
Published: (2025)
by: Gao, Chen, et al.
Published: (2025)
VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
by: Lin, Sihao, et al.
Published: (2025)
by: Lin, Sihao, et al.
Published: (2025)
Towards Precise Intent-Aligned VLA Aerial Navigation via Expert-Guided GRPO
by: Chen, Tianyang, et al.
Published: (2026)
by: Chen, Tianyang, et al.
Published: (2026)
VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory
by: Wang, Shaoan, et al.
Published: (2026)
by: Wang, Shaoan, et al.
Published: (2026)
ReWorld: Multi-Dimensional Reward Modeling for Embodied World Models
by: Peng, Baorui, et al.
Published: (2026)
by: Peng, Baorui, et al.
Published: (2026)
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
by: Zhou, Gengze, et al.
Published: (2024)
by: Zhou, Gengze, et al.
Published: (2024)
Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments
by: Li, Zerui, et al.
Published: (2025)
by: Li, Zerui, et al.
Published: (2025)
Pediatric Chronic Monteggia Fractures: Insights From a Comprehensive Review
by: Gengze Li, et al.
Published: (2025)
by: Gengze Li, et al.
Published: (2025)
LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion
by: Lyu, Jiangran, et al.
Published: (2026)
by: Lyu, Jiangran, et al.
Published: (2026)
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
by: Zhou, Gengze, et al.
Published: (2024)
by: Zhou, Gengze, et al.
Published: (2024)
ABot-N0: Technical Report on the VLA Foundation Model for Versatile Embodied Navigation
by: Chu, Zedong, et al.
Published: (2026)
by: Chu, Zedong, et al.
Published: (2026)
Lifelong Embodied Navigation Learning
by: Wang, Xudong, et al.
Published: (2026)
by: Wang, Xudong, et al.
Published: (2026)
Learning Goal-Oriented Vision-and-Language Navigation with Self-Improving Demonstrations at Scale
by: Li, Songze, et al.
Published: (2025)
by: Li, Songze, et al.
Published: (2025)
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
by: Zhang, Wenyao, et al.
Published: (2025)
by: Zhang, Wenyao, et al.
Published: (2025)
PathPainter: Transferring the Generalization Ability of Image Generation Models to Embodied Navigation
by: Wang, Yijin, et al.
Published: (2026)
by: Wang, Yijin, et al.
Published: (2026)
MetaUrban: An Embodied AI Simulation Platform for Urban Micromobility
by: Wu, Wayne, et al.
Published: (2024)
by: Wu, Wayne, et al.
Published: (2024)
Agentic Self-Evolutionary Replanning for Embodied Navigation
by: Li, Guoliang, et al.
Published: (2026)
by: Li, Guoliang, et al.
Published: (2026)
SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation
by: Zhang, Junjie, et al.
Published: (2024)
by: Zhang, Junjie, et al.
Published: (2024)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
by: Li, Manling, et al.
Published: (2024)
by: Li, Manling, et al.
Published: (2024)
NavBench: Probing Multimodal Large Language Models for Embodied Navigation
by: Qiao, Yanyuan, et al.
Published: (2025)
by: Qiao, Yanyuan, et al.
Published: (2025)
Universal Actions for Enhanced Embodied Foundation Models
by: Zheng, Jinliang, et al.
Published: (2025)
by: Zheng, Jinliang, et al.
Published: (2025)
Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
Towards Efficient LLM Grounding for Embodied Multi-Agent Collaboration
by: Zhang, Yang, et al.
Published: (2024)
by: Zhang, Yang, et al.
Published: (2024)
Rate-Distortion-Cognition Controllable Versatile Neural Image Compression
by: Liu, Jinming, et al.
Published: (2024)
by: Liu, Jinming, et al.
Published: (2024)
Joint Auction in the Online Advertising Market
by: Zhang, Zhen, et al.
Published: (2024)
by: Zhang, Zhen, et al.
Published: (2024)
HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions
by: Dong, Yifei, et al.
Published: (2025)
by: Dong, Yifei, et al.
Published: (2025)
Progress in Synthesis of CO 2 ‐Based Polycarbonate Diols and Their Application in Preparation of Different Types of Polyurethane
by: Qi Yan, et al.
Published: (2026)
by: Qi Yan, et al.
Published: (2026)
Embodied World Models Emerge from Navigational Task in Open-Ended Environments
by: Jin, Li, et al.
Published: (2025)
by: Jin, Li, et al.
Published: (2025)
Do News and Social Media Tell the Same Story? Constructing and Comparing Sentiment Spillover Networks
by: Wu, Fan, et al.
Published: (2026)
by: Wu, Fan, et al.
Published: (2026)
An Initial Investigation of Neural Replay Simulator for Over-the-Air Adversarial Perturbations to Automatic Speaker Verification
by: Li, Jiaqi, et al.
Published: (2023)
by: Li, Jiaqi, et al.
Published: (2023)
Similar Items
-
TrackVLA: Embodied Visual Tracking in the Wild
by: Wang, Shaoan, et al.
Published: (2025) -
TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
by: Liu, Jiahang, et al.
Published: (2025) -
NavGSim: High-Fidelity Gaussian Splatting Simulator for Large-Scale Navigation
by: Liu, Jiahang, et al.
Published: (2026) -
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
by: Zhang, Jiazhao, et al.
Published: (2024) -
UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
by: Li, Anqi, et al.
Published: (2025)