Towards Open Environments and Instructions: General Vision-Language Navigation via Fast-Slow Interactive Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yang, Wu, Aming, Zhang, Zihao, Han, Yahong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Novel Class Discovery for Point Cloud Segmentation via Joint Learning of Causal Representation and Reasoning
von: Li, Yang, et al.
Veröffentlicht: (2025)
von: Li, Yang, et al.
Veröffentlicht: (2025)
Simulating Distribution Dynamics: Liquid Temporal Feature Evolution for Single-Domain Generalized Object Detection
von: Zhang, Zihao, et al.
Veröffentlicht: (2025)
von: Zhang, Zihao, et al.
Veröffentlicht: (2025)
Style Evolving along Chain-of-Thought for Unknown-Domain Object Detection
von: Zhang, Zihao, et al.
Veröffentlicht: (2025)
von: Zhang, Zihao, et al.
Veröffentlicht: (2025)
Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation
von: Zhang, Xitie, et al.
Veröffentlicht: (2026)
von: Zhang, Xitie, et al.
Veröffentlicht: (2026)
Continual Adaptation: Environment-Conditional Parameter Generation for Object Detection in Dynamic Scenarios
von: Li, Deng, et al.
Veröffentlicht: (2025)
von: Li, Deng, et al.
Veröffentlicht: (2025)
Prompt-Driven Dynamic Object-Centric Learning for Single Domain Generalization
von: Li, Deng, et al.
Veröffentlicht: (2024)
von: Li, Deng, et al.
Veröffentlicht: (2024)
Fast-Slow Test-Time Adaptation for Online Vision-and-Language Navigation
von: Gao, Junyu, et al.
Veröffentlicht: (2023)
von: Gao, Junyu, et al.
Veröffentlicht: (2023)
StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
von: Wei, Meng, et al.
Veröffentlicht: (2025)
von: Wei, Meng, et al.
Veröffentlicht: (2025)
Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
SlowFastVAD: Video Anomaly Detection via Integrating Simple Detector and RAG-Enhanced Vision-Language Model
von: Ding, Zongcan, et al.
Veröffentlicht: (2025)
von: Ding, Zongcan, et al.
Veröffentlicht: (2025)
Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs
von: Qiao, Yanyuan, et al.
Veröffentlicht: (2024)
von: Qiao, Yanyuan, et al.
Veröffentlicht: (2024)
Generalizing to Out-of-Sample Degradations via Model Reprogramming
von: Jiang, Runhua, et al.
Veröffentlicht: (2024)
von: Jiang, Runhua, et al.
Veröffentlicht: (2024)
Volumetric Environment Representation for Vision-Language Navigation
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
von: Xiao, Wenyi, et al.
Veröffentlicht: (2025)
von: Xiao, Wenyi, et al.
Veröffentlicht: (2025)
A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
von: Li, Chenxuan, et al.
Veröffentlicht: (2024)
von: Li, Chenxuan, et al.
Veröffentlicht: (2024)
World-Consistent Data Generation for Vision-and-Language Navigation
von: Zhong, Yu, et al.
Veröffentlicht: (2024)
von: Zhong, Yu, et al.
Veröffentlicht: (2024)
ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation
von: Xue, Wei, et al.
Veröffentlicht: (2026)
von: Xue, Wei, et al.
Veröffentlicht: (2026)
Visual Consensus Prompting for Co-Salient Object Detection
von: Wang, Jie, et al.
Veröffentlicht: (2025)
von: Wang, Jie, et al.
Veröffentlicht: (2025)
Stop Wandering: Efficient Vision-Language Navigation via Metacognitive Reasoning
von: Li, Xueying, et al.
Veröffentlicht: (2026)
von: Li, Xueying, et al.
Veröffentlicht: (2026)
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
von: Yang, Jie, et al.
Veröffentlicht: (2025)
von: Yang, Jie, et al.
Veröffentlicht: (2025)
Learning to Think Fast and Slow for Visual Language Models
von: Lin, Chenyu, et al.
Veröffentlicht: (2025)
von: Lin, Chenyu, et al.
Veröffentlicht: (2025)
AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning
von: Chen, Weixing, et al.
Veröffentlicht: (2025)
von: Chen, Weixing, et al.
Veröffentlicht: (2025)
PaveBench: A Versatile Benchmark for Pavement Distress Perception and Interactive Vision-Language Analysis
von: Li, Dexiang, et al.
Veröffentlicht: (2026)
von: Li, Dexiang, et al.
Veröffentlicht: (2026)
Fine-Grained Instruction-Guided Graph Reasoning for Vision-and-Language Navigation
von: Liu, Yaohua, et al.
Veröffentlicht: (2025)
von: Liu, Yaohua, et al.
Veröffentlicht: (2025)
Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations
von: Cui, Yibo, et al.
Veröffentlicht: (2025)
von: Cui, Yibo, et al.
Veröffentlicht: (2025)
Interactive Continual Learning: Fast and Slow Thinking
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
Towards Zero-Shot Annotation of the Built Environment with Vision-Language Models (Vision Paper)
von: Han, Bin, et al.
Veröffentlicht: (2024)
von: Han, Bin, et al.
Veröffentlicht: (2024)
SlowFast-SCI: Slow-Fast Deep Unfolding Learning for Spectral Compressive Imaging
von: Zeng, Haijin, et al.
Veröffentlicht: (2025)
von: Zeng, Haijin, et al.
Veröffentlicht: (2025)
Navigation Instruction Generation with BEV Perception and Large Language Models
von: Fan, Sheng, et al.
Veröffentlicht: (2024)
von: Fan, Sheng, et al.
Veröffentlicht: (2024)
Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method
von: Song, Xinshuai, et al.
Veröffentlicht: (2024)
von: Song, Xinshuai, et al.
Veröffentlicht: (2024)
Towards Realistic UAV Vision-Language Navigation: Platform, Benchmark, and Methodology
von: Wang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Wang, Xiangyu, et al.
Veröffentlicht: (2024)
Hierarchical Spatial Proximity Reasoning for Vision-and-Language Navigation
von: Xu, Ming, et al.
Veröffentlicht: (2024)
von: Xu, Ming, et al.
Veröffentlicht: (2024)
Emotional Theory of Mind: Bridging Fast Visual Processing with Slow Linguistic Reasoning
von: Etesam, Yasaman, et al.
Veröffentlicht: (2023)
von: Etesam, Yasaman, et al.
Veröffentlicht: (2023)
MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
von: Chu, Xiangxiang, et al.
Veröffentlicht: (2023)
von: Chu, Xiangxiang, et al.
Veröffentlicht: (2023)
FASTopoWM: Fast-Slow Lane Segment Topology Reasoning with Latent World Models
von: Yang, Yiming, et al.
Veröffentlicht: (2025)
von: Yang, Yiming, et al.
Veröffentlicht: (2025)
PROGRESSLM: Towards Progress Reasoning in Vision-Language Models
von: Zhang, Jianshu, et al.
Veröffentlicht: (2026)
von: Zhang, Jianshu, et al.
Veröffentlicht: (2026)
Towards Robust and Fair Vision Learning in Open-World Environments
von: Truong, Thanh-Dat
Veröffentlicht: (2024)
von: Truong, Thanh-Dat
Veröffentlicht: (2024)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
Fast-SmartWay: Panoramic-Free End-to-End Zero-Shot Vision-and-Language Navigation
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive Interaction
von: Li, Yuanbo, et al.
Veröffentlicht: (2026)
von: Li, Yuanbo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Novel Class Discovery for Point Cloud Segmentation via Joint Learning of Causal Representation and Reasoning
von: Li, Yang, et al.
Veröffentlicht: (2025) -
Simulating Distribution Dynamics: Liquid Temporal Feature Evolution for Single-Domain Generalized Object Detection
von: Zhang, Zihao, et al.
Veröffentlicht: (2025) -
Style Evolving along Chain-of-Thought for Unknown-Domain Object Detection
von: Zhang, Zihao, et al.
Veröffentlicht: (2025) -
Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation
von: Zhang, Xitie, et al.
Veröffentlicht: (2026) -
Continual Adaptation: Environment-Conditional Parameter Generation for Object Detection in Dynamic Scenarios
von: Li, Deng, et al.
Veröffentlicht: (2025)