Let's Reward Step-by-Step: Step-Aware Contrastive Alignment for Vision-Language Navigation in Continuous Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Haoyuan, Liu, Rui, Fan, Hehe, Yang, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
by: Li, Haoyuan, et al.
Published: (2025)
by: Li, Haoyuan, et al.
Published: (2025)
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
by: Zhang, Jiazhao, et al.
Published: (2024)
by: Zhang, Jiazhao, et al.
Published: (2024)
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
by: Xu, Guowei, et al.
Published: (2024)
by: Xu, Guowei, et al.
Published: (2024)
Three-Step Nav: A Hierarchical Global-Local Planner for Zero-Shot Vision-and-Language Navigation
by: Zheng, Wanrong, et al.
Published: (2026)
by: Zheng, Wanrong, et al.
Published: (2026)
Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
by: Chen, Kehan, et al.
Published: (2024)
by: Chen, Kehan, et al.
Published: (2024)
How Good are Foundation Models in Step-by-Step Embodied Reasoning?
by: Dissanayake, Dinura, et al.
Published: (2025)
by: Dissanayake, Dinura, et al.
Published: (2025)
Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
by: Lu, Jinghui, et al.
Published: (2026)
by: Lu, Jinghui, et al.
Published: (2026)
$π$-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs
by: Wang, Siting, et al.
Published: (2026)
by: Wang, Siting, et al.
Published: (2026)
DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding
by: Ishaq, Ayesha, et al.
Published: (2025)
by: Ishaq, Ayesha, et al.
Published: (2025)
Think Step by Step: Chain-of-Gesture Prompting for Error Detection in Robotic Surgical Videos
by: Shao, Zhimin, et al.
Published: (2024)
by: Shao, Zhimin, et al.
Published: (2024)
Lookahead Exploration with Neural Radiance Representation for Continuous Vision-Language Navigation
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments
by: An, Dong, et al.
Published: (2023)
by: An, Dong, et al.
Published: (2023)
Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs
by: Qiao, Yanyuan, et al.
Published: (2024)
by: Qiao, Yanyuan, et al.
Published: (2024)
Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments
by: Hong, Haodong, et al.
Published: (2024)
by: Hong, Haodong, et al.
Published: (2024)
A Step Toward World Models: A Survey on Robotic Manipulation
by: Zhang, Peng-Fei, et al.
Published: (2025)
by: Zhang, Peng-Fei, et al.
Published: (2025)
ST-Booster: An Iterative SpatioTemporal Perception Booster for Vision-and-Language Navigation in Continuous Environments
by: Yue, Lu, et al.
Published: (2025)
by: Yue, Lu, et al.
Published: (2025)
View Invariant Learning for Vision-Language Navigation in Continuous Environments
by: Sun, Josh Qixuan, et al.
Published: (2025)
by: Sun, Josh Qixuan, et al.
Published: (2025)
AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
by: Guo, Wenxuan, et al.
Published: (2026)
by: Guo, Wenxuan, et al.
Published: (2026)
Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments
by: Li, Zerui, et al.
Published: (2025)
by: Li, Zerui, et al.
Published: (2025)
Need for Speed: Zero-Shot Depth Completion with Single-Step Diffusion
by: Gregorek, Jakub, et al.
Published: (2026)
by: Gregorek, Jakub, et al.
Published: (2026)
TagaVLM: Topology-Aware Global Action Reasoning for Vision-Language Navigation
by: Liu, Jiaxing, et al.
Published: (2026)
by: Liu, Jiaxing, et al.
Published: (2026)
ReAlign: Text-to-Motion Generation via Step-Aware Reward-Guided Alignment
by: Weng, Wanjiang, et al.
Published: (2025)
by: Weng, Wanjiang, et al.
Published: (2025)
CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step
by: Liu, Zheyuan, et al.
Published: (2025)
by: Liu, Zheyuan, et al.
Published: (2025)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Real-Time Metric-Semantic Mapping for Autonomous Navigation in Outdoor Environments
by: Jiao, Jianhao, et al.
Published: (2024)
by: Jiao, Jianhao, et al.
Published: (2024)
Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning
by: Bardhan, Jai, et al.
Published: (2026)
by: Bardhan, Jai, et al.
Published: (2026)
ReAlign: Bilingual Text-to-Motion Generation via Step-Aware Reward-Guided Alignment
by: Weng, Wanjiang, et al.
Published: (2025)
by: Weng, Wanjiang, et al.
Published: (2025)
Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
by: Chen, Honghao, et al.
Published: (2025)
by: Chen, Honghao, et al.
Published: (2025)
Volumetric Environment Representation for Vision-Language Navigation
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming
by: Wang, Shuo, et al.
Published: (2025)
by: Wang, Shuo, et al.
Published: (2025)
AdaVLN: Towards Visual Language Navigation in Continuous Indoor Environments with Moving Humans
by: Loh, Dillon, et al.
Published: (2024)
by: Loh, Dillon, et al.
Published: (2024)
FlowSSC: Universal Generative Monocular Semantic Scene Completion via One-Step Latent Diffusion
by: Xi, Zichen, et al.
Published: (2026)
by: Xi, Zichen, et al.
Published: (2026)
Sim-to-Real Transfer via 3D Feature Fields for Vision-and-Language Navigation
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
Stop Wandering: Efficient Vision-Language Navigation via Metacognitive Reasoning
by: Li, Xueying, et al.
Published: (2026)
by: Li, Xueying, et al.
Published: (2026)
Slow Perception: Let's Perceive Geometric Figures Step-by-step
by: Wei, Haoran, et al.
Published: (2024)
by: Wei, Haoran, et al.
Published: (2024)
Enhancing Vision-Language Navigation with Multimodal Event Knowledge from Real-World Indoor Tour Videos
by: Xu, Haoxuan, et al.
Published: (2026)
by: Xu, Haoxuan, et al.
Published: (2026)
Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning
by: Qi, Xiuxiu, et al.
Published: (2025)
by: Qi, Xiuxiu, et al.
Published: (2025)
MeanFuser: Fast One-Step Multi-Modal Trajectory Generation and Adaptive Reconstruction via MeanFlow for End-to-End Autonomous Driving
by: Wang, Junli, et al.
Published: (2026)
by: Wang, Junli, et al.
Published: (2026)
Towards Realistic UAV Vision-Language Navigation: Platform, Benchmark, and Methodology
by: Wang, Xiangyu, et al.
Published: (2024)
by: Wang, Xiangyu, et al.
Published: (2024)
Step by Step Network
by: Han, Dongchen, et al.
Published: (2025)
by: Han, Dongchen, et al.
Published: (2025)
Similar Items
-
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
by: Li, Haoyuan, et al.
Published: (2025) -
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
by: Zhang, Jiazhao, et al.
Published: (2024) -
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
by: Xu, Guowei, et al.
Published: (2024) -
Three-Step Nav: A Hierarchical Global-Local Planner for Zero-Shot Vision-and-Language Navigation
by: Zheng, Wanrong, et al.
Published: (2026) -
Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
by: Chen, Kehan, et al.
Published: (2024)