AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation
Fuente:
arXiv
Saved in:
| Main Authors: | Ding, Xin, Wei, Jianyu, Yang, Yifan, Jiang, Shiqi, Zhang, Qianxi, Wu, Hao, Jia, Fucheng, Mi, Liang, Yan, Yuxuan, Wang, Weijun, Liu, Yunxin, Chen, Zhibo, Cao, Ting |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding
by: Zheng, Yikai, et al.
Published: (2026)
by: Zheng, Yikai, et al.
Published: (2026)
Making Every Frame Matter: Continuous Activity Recognition in Streaming Video via Adaptive Video Context Modeling
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents
by: Ding, Xin, et al.
Published: (2026)
by: Ding, Xin, et al.
Published: (2026)
Scaling Up On-Device LLMs via Active-Weight Swapping Between DRAM and Flash
by: Jia, Fucheng, et al.
Published: (2025)
by: Jia, Fucheng, et al.
Published: (2025)
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models
by: Huang, Mingzhe, et al.
Published: (2026)
by: Huang, Mingzhe, et al.
Published: (2026)
Nav-R1: Reasoning and Navigation in Embodied Scenes
by: Liu, Qingxiang, et al.
Published: (2025)
by: Liu, Qingxiang, et al.
Published: (2025)
Efficient Remote KV Cache Reuse with GPU-native Video Codec
by: Mi, Liang, et al.
Published: (2026)
by: Mi, Liang, et al.
Published: (2026)
EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents
by: Ju, Ruofei, et al.
Published: (2026)
by: Ju, Ruofei, et al.
Published: (2026)
StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition
by: Ding, Xin, et al.
Published: (2025)
by: Ding, Xin, et al.
Published: (2025)
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
by: Li, Xiangyu, et al.
Published: (2025)
by: Li, Xiangyu, et al.
Published: (2025)
Dissecting Bit-Level Scaling Laws in Quantizing Vision Generative Models
by: Ding, Xin, et al.
Published: (2025)
by: Ding, Xin, et al.
Published: (2025)
OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism
by: Li, Xiangyu, et al.
Published: (2026)
by: Li, Xiangyu, et al.
Published: (2026)
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
by: Zhou, Gengze, et al.
Published: (2024)
by: Zhou, Gengze, et al.
Published: (2024)
Hydra-Nav: Object Navigation via Adaptive Dual-Process Reasoning
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning
by: Lin, Bingqian, et al.
Published: (2024)
by: Lin, Bingqian, et al.
Published: (2024)
CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
by: Hao, Haihong, et al.
Published: (2025)
by: Hao, Haihong, et al.
Published: (2025)
VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory
by: Wang, Shaoan, et al.
Published: (2026)
by: Wang, Shaoan, et al.
Published: (2026)
EvolveNav: Empowering LLM-Based Vision-Language Navigation via Self-Improving Embodied Reasoning
by: Lin, Bingqian, et al.
Published: (2025)
by: Lin, Bingqian, et al.
Published: (2025)
VL-Nav: A Neuro-Symbolic Approach for Reasoning-based Vision-Language Navigation
by: Du, Yi, et al.
Published: (2025)
by: Du, Yi, et al.
Published: (2025)
SoraNav: Adaptive UAV Task-Centric Navigation via Zeroshot VLM Reasoning
by: Song, Hongyu, et al.
Published: (2025)
by: Song, Hongyu, et al.
Published: (2025)
Empowering In-Browser Deep Learning Inference on Edge Devices with Just-in-Time Kernel Optimizations
by: Jia, Fucheng, et al.
Published: (2023)
by: Jia, Fucheng, et al.
Published: (2023)
IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
by: Ding, Liang
Published: (2026)
by: Ding, Liang
Published: (2026)
CorNav: Autonomous Agent with Self-Corrected Planning for Zero-Shot Vision-and-Language Navigation
by: Liang, Xiwen, et al.
Published: (2023)
by: Liang, Xiwen, et al.
Published: (2023)
Uncertainty-Aware Gaussian Map for Vision-Language Navigation
by: Gao, Jianzhe, et al.
Published: (2026)
by: Gao, Jianzhe, et al.
Published: (2026)
AVA: Towards Agentic Video Analytics with Vision Language Models
by: Yan, Yuxuan, et al.
Published: (2025)
by: Yan, Yuxuan, et al.
Published: (2025)
AdaKernel: Learning Adaptive Kernel Parameters for Spatiotemporal Graph Neural Networks
by: Zhang, Zhongyue, et al.
Published: (2026)
by: Zhang, Zhongyue, et al.
Published: (2026)
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
by: Yu, Zhuoyuan, et al.
Published: (2025)
by: Yu, Zhuoyuan, et al.
Published: (2025)
AdaPRL: Adaptive Pairwise Regression Learning with Uncertainty Estimation for Universal Regression Tasks
by: Liang, Fuhang, et al.
Published: (2025)
by: Liang, Fuhang, et al.
Published: (2025)
OLiVia-Nav: An Online Lifelong Vision Language Approach for Mobile Robot Social Navigation
by: Narasimhan, Siddarth, et al.
Published: (2024)
by: Narasimhan, Siddarth, et al.
Published: (2024)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
by: Zhang, Minxin, et al.
Published: (2025)
by: Zhang, Minxin, et al.
Published: (2025)
Region-based Content Enhancement for Efficient Video Analytics at the Edge
by: Wang, Weijun, et al.
Published: (2024)
by: Wang, Weijun, et al.
Published: (2024)
BiSwift: Bandwidth Orchestrator for Multi-Stream Video Analytics on Edge
by: Sun, Lin, et al.
Published: (2023)
by: Sun, Lin, et al.
Published: (2023)
Co-NavGPT: Multi-Robot Cooperative Visual Semantic Navigation Using Vision Language Models
by: Yu, Bangguo, et al.
Published: (2023)
by: Yu, Bangguo, et al.
Published: (2023)
MAER-Nav: Bidirectional Motion Learning Through Mirror-Augmented Experience Replay for Robot Navigation
by: Wang, Shanze, et al.
Published: (2025)
by: Wang, Shanze, et al.
Published: (2025)
SPAN-Nav: Generalized Spatial Awareness for Versatile Vision-Language Navigation
by: Liu, Jiahang, et al.
Published: (2026)
by: Liu, Jiahang, et al.
Published: (2026)
NavHint: Vision and Language Navigation Agent with a Hint Generator
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
FeudalNav: A Simple Framework for Visual Navigation
by: Johnson, Faith, et al.
Published: (2026)
by: Johnson, Faith, et al.
Published: (2026)
AdaContour: Adaptive Contour Descriptor with Hierarchical Representation
by: Ding, Tianyu, et al.
Published: (2024)
by: Ding, Tianyu, et al.
Published: (2024)
MOFM-Nav: On-Manifold Ordering-Flexible Multi-Robot Navigation
by: Hu, Bin-Bin, et al.
Published: (2025)
by: Hu, Bin-Bin, et al.
Published: (2025)
Similar Items
-
Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding
by: Zheng, Yikai, et al.
Published: (2026) -
Making Every Frame Matter: Continuous Activity Recognition in Streaming Video via Adaptive Video Context Modeling
by: Wu, Hao, et al.
Published: (2024) -
MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents
by: Ding, Xin, et al.
Published: (2026) -
Scaling Up On-Device LLMs via Active-Weight Swapping Between DRAM and Flash
by: Jia, Fucheng, et al.
Published: (2025) -
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models
by: Huang, Mingzhe, et al.
Published: (2026)