RATE-Nav: Region-Aware Termination Enhancement for Zero-shot Object Navigation with Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Junjie, Zhang, Nan, Qu, Xiaoyang, Lu, Kai, Li, Guokuan, Wan, Jiguang, Wang, Jianzong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
di: Zhang, Bin, et al.
Pubblicazione: (2025)
di: Zhang, Bin, et al.
Pubblicazione: (2025)
Triage: Hierarchical Visual Budgeting for Efficient Video Reasoning in Vision-Language Models
di: Wang, Anmin, et al.
Pubblicazione: (2026)
di: Wang, Anmin, et al.
Pubblicazione: (2026)
RUNA: Object-level Out-of-Distribution Detection via Regional Uncertainty Alignment of Multimodal Representations
di: Zhang, Bin, et al.
Pubblicazione: (2025)
di: Zhang, Bin, et al.
Pubblicazione: (2025)
Vista: Scene-Aware Optimization for Streaming Video Question Answering under Post-Hoc Queries
di: Lu, Haocheng, et al.
Pubblicazione: (2026)
di: Lu, Haocheng, et al.
Pubblicazione: (2026)
PRENet: A Plane-Fit Redundancy Encoding Point Cloud Sequence Network for Real-Time 3D Action Recognition
di: He, Shenglin, et al.
Pubblicazione: (2024)
di: He, Shenglin, et al.
Pubblicazione: (2024)
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
di: Tao, Wei, et al.
Pubblicazione: (2026)
di: Tao, Wei, et al.
Pubblicazione: (2026)
Value-Driven Mixed-Precision Quantization for Patch-Based Inference on Microcontrollers
di: Tao, Wei, et al.
Pubblicazione: (2024)
di: Tao, Wei, et al.
Pubblicazione: (2024)
BAGNet: A Boundary-Aware Graph Attention Network for 3D Point Cloud Semantic Segmentation
di: Tao, Wei, et al.
Pubblicazione: (2025)
di: Tao, Wei, et al.
Pubblicazione: (2025)
MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts
di: Tao, Wei, et al.
Pubblicazione: (2025)
di: Tao, Wei, et al.
Pubblicazione: (2025)
VoroNav: Voronoi-based Zero-shot Object Navigation with Large Language Model
di: Wu, Pengying, et al.
Pubblicazione: (2024)
di: Wu, Pengying, et al.
Pubblicazione: (2024)
SR-Nav: Spatial Relationships Matter for Zero-shot Object Goal Navigation
di: Fang, Leyuan, et al.
Pubblicazione: (2026)
di: Fang, Leyuan, et al.
Pubblicazione: (2026)
Enhancing Multi-Agent Systems via Reinforcement Learning with LLM-based Planner and Graph-based Policy
di: Jia, Ziqi, et al.
Pubblicazione: (2025)
di: Jia, Ziqi, et al.
Pubblicazione: (2025)
MADLLM: Multivariate Anomaly Detection via Pre-trained LLMs
di: Tao, Wei, et al.
Pubblicazione: (2025)
di: Tao, Wei, et al.
Pubblicazione: (2025)
SG-Nav: Online 3D Scene Graph Prompting for LLM-based Zero-shot Object Navigation
di: Yin, Hang, et al.
Pubblicazione: (2024)
di: Yin, Hang, et al.
Pubblicazione: (2024)
VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success
di: Liu, Chuhang, et al.
Pubblicazione: (2026)
di: Liu, Chuhang, et al.
Pubblicazione: (2026)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
di: Zhang, Jiwen, et al.
Pubblicazione: (2026)
di: Zhang, Jiwen, et al.
Pubblicazione: (2026)
ConsistNav: Closing the Action Consistency Gap in Zero-Shot Object Navigation with Semantic Executive Control
di: Wang, Haosen, et al.
Pubblicazione: (2026)
di: Wang, Haosen, et al.
Pubblicazione: (2026)
From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration
di: Shi, Jiaqi, et al.
Pubblicazione: (2026)
di: Shi, Jiaqi, et al.
Pubblicazione: (2026)
AERR-Nav: Adaptive Exploration-Recovery-Reminiscing Strategy for Zero-Shot Object Navigation
di: Huang, Jingzhi, et al.
Pubblicazione: (2026)
di: Huang, Jingzhi, et al.
Pubblicazione: (2026)
TopV-Nav: Unlocking the Top-View Spatial Reasoning Potential of MLLM for Zero-shot Object Navigation
di: Zhong, Linqing, et al.
Pubblicazione: (2024)
di: Zhong, Linqing, et al.
Pubblicazione: (2024)
P2DNav: Panorama-to-Downview Reasoning for Zero-shot Vision-and-Language Navigation
di: Sheng, Kai, et al.
Pubblicazione: (2026)
di: Sheng, Kai, et al.
Pubblicazione: (2026)
Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs
di: Qiao, Yanyuan, et al.
Pubblicazione: (2024)
di: Qiao, Yanyuan, et al.
Pubblicazione: (2024)
LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs
di: Luo, Kun, et al.
Pubblicazione: (2026)
di: Luo, Kun, et al.
Pubblicazione: (2026)
Hierarchical-Task-Aware Multi-modal Mixture of Incremental LoRA Experts for Embodied Continual Learning
di: Jia, Ziqi, et al.
Pubblicazione: (2025)
di: Jia, Ziqi, et al.
Pubblicazione: (2025)
InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment
di: Long, Yuxing, et al.
Pubblicazione: (2024)
di: Long, Yuxing, et al.
Pubblicazione: (2024)
FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation
di: Shao, Dian, et al.
Pubblicazione: (2026)
di: Shao, Dian, et al.
Pubblicazione: (2026)
Zero-shot Generalizable Incremental Learning for Vision-Language Object Detection
di: Deng, Jieren, et al.
Pubblicazione: (2024)
di: Deng, Jieren, et al.
Pubblicazione: (2024)
Three-Step Nav: A Hierarchical Global-Local Planner for Zero-Shot Vision-and-Language Navigation
di: Zheng, Wanrong, et al.
Pubblicazione: (2026)
di: Zheng, Wanrong, et al.
Pubblicazione: (2026)
Federated Domain Generalization with Domain-specific Soft Prompts Generation
di: Wu, Jianhan, et al.
Pubblicazione: (2025)
di: Wu, Jianhan, et al.
Pubblicazione: (2025)
Task-Specific Zero-shot Quantization-Aware Training for Object Detection
di: Li, Changhao, et al.
Pubblicazione: (2025)
di: Li, Changhao, et al.
Pubblicazione: (2025)
MIRRORTALK: Forging Personalized Avatars Via Disentangled Style and Hierarchical Motion Control
di: Lu, Renjie, et al.
Pubblicazione: (2026)
di: Lu, Renjie, et al.
Pubblicazione: (2026)
DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement
di: Lu, Renjie, et al.
Pubblicazione: (2026)
di: Lu, Renjie, et al.
Pubblicazione: (2026)
CorNav: Autonomous Agent with Self-Corrected Planning for Zero-Shot Vision-and-Language Navigation
di: Liang, Xiwen, et al.
Pubblicazione: (2023)
di: Liang, Xiwen, et al.
Pubblicazione: (2023)
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
di: Shi, Jiaqi, et al.
Pubblicazione: (2026)
di: Shi, Jiaqi, et al.
Pubblicazione: (2026)
PanoNav: Mapless Zero-Shot Object Navigation with Panoramic Scene Parsing and Dynamic Memory
di: Jin, Qunchao, et al.
Pubblicazione: (2025)
di: Jin, Qunchao, et al.
Pubblicazione: (2025)
ReMemNav: A Rethinking and Memory-Augmented Framework for Zero-Shot Object Navigation
di: Wu, Feng, et al.
Pubblicazione: (2026)
di: Wu, Feng, et al.
Pubblicazione: (2026)
FOM-Nav: Frontier-Object Maps for Object Goal Navigation
di: Chabal, Thomas, et al.
Pubblicazione: (2025)
di: Chabal, Thomas, et al.
Pubblicazione: (2025)
VL-Nav: A Neuro-Symbolic Approach for Reasoning-based Vision-Language Navigation
di: Du, Yi, et al.
Pubblicazione: (2025)
di: Du, Yi, et al.
Pubblicazione: (2025)
DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
di: Wang, Yunheng, et al.
Pubblicazione: (2025)
di: Wang, Yunheng, et al.
Pubblicazione: (2025)
DOPE: Dual Object Perception-Enhancement Network for Vision-and-Language Navigation
di: Yu, Yinfeng, et al.
Pubblicazione: (2025)
di: Yu, Yinfeng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
di: Zhang, Bin, et al.
Pubblicazione: (2025) -
Triage: Hierarchical Visual Budgeting for Efficient Video Reasoning in Vision-Language Models
di: Wang, Anmin, et al.
Pubblicazione: (2026) -
RUNA: Object-level Out-of-Distribution Detection via Regional Uncertainty Alignment of Multimodal Representations
di: Zhang, Bin, et al.
Pubblicazione: (2025) -
Vista: Scene-Aware Optimization for Streaming Video Question Answering under Post-Hoc Queries
di: Lu, Haocheng, et al.
Pubblicazione: (2026) -
PRENet: A Plane-Fit Redundancy Encoding Point Cloud Sequence Network for Real-Time 3D Action Recognition
di: He, Shenglin, et al.
Pubblicazione: (2024)