Seeing Farther and Smarter: Value-Guided Multi-Path Reflection for VLM Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Yanting, Gao, Shenyuan, Bu, Qingwen, Chen, Li, Metaxas, Dimitris N. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
by: Xu, Kechun, et al.
Published: (2025)
by: Xu, Kechun, et al.
Published: (2025)
StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
by: Liu, Mingyu, et al.
Published: (2025)
by: Liu, Mingyu, et al.
Published: (2025)
VLM-Guided Visual Place Recognition for Planet-Scale Geo-Localization
by: Waheed, Sania, et al.
Published: (2025)
by: Waheed, Sania, et al.
Published: (2025)
VDT-Auto: End-to-end Autonomous Driving with VLM-Guided Diffusion Transformers
by: Guo, Ziang, et al.
Published: (2025)
by: Guo, Ziang, et al.
Published: (2025)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
ReSim: Reliable World Simulation for Autonomous Driving
by: Yang, Jiazhi, et al.
Published: (2025)
by: Yang, Jiazhi, et al.
Published: (2025)
NaturalVLM: Leveraging Fine-grained Natural Language for Affordance-Guided Visual Manipulation
by: Xu, Ran, et al.
Published: (2024)
by: Xu, Ran, et al.
Published: (2024)
Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation
by: Gu, Difei, et al.
Published: (2025)
by: Gu, Difei, et al.
Published: (2025)
MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model
by: Dao, Quan, et al.
Published: (2026)
by: Dao, Quan, et al.
Published: (2026)
DreamAvoid: Critical-Phase Test-Time Dreaming to Avoid Failures in VLA Policies
by: Fan, Xianzhe, et al.
Published: (2026)
by: Fan, Xianzhe, et al.
Published: (2026)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
by: Li, Rong, et al.
Published: (2024)
by: Li, Rong, et al.
Published: (2024)
Enhancing End-to-End Autonomous Driving with Risk Semantic Distillaion from VLM
by: Qin, Jack, et al.
Published: (2025)
by: Qin, Jack, et al.
Published: (2025)
Score-Guided Diffusion for 3D Human Recovery
by: Stathopoulos, Anastasis, et al.
Published: (2024)
by: Stathopoulos, Anastasis, et al.
Published: (2024)
DeltaFlow: An Efficient Multi-frame Scene Flow Estimation Method
by: Zhang, Qingwen, et al.
Published: (2025)
by: Zhang, Qingwen, et al.
Published: (2025)
SPARK: Sim-ready Part-level Articulated Reconstruction with VLM Knowledge
by: He, Yumeng, et al.
Published: (2025)
by: He, Yumeng, et al.
Published: (2025)
VLMFusionOcc3D: VLM Assisted Multi-Modal 3D Semantic Occupancy Prediction
by: Doruk, A. Enes, et al.
Published: (2026)
by: Doruk, A. Enes, et al.
Published: (2026)
SeFlow: A Self-Supervised Scene Flow Method in Autonomous Driving
by: Zhang, Qingwen, et al.
Published: (2024)
by: Zhang, Qingwen, et al.
Published: (2024)
DeFlow: Decoder of Scene Flow Network in Autonomous Driving
by: Zhang, Qingwen, et al.
Published: (2024)
by: Zhang, Qingwen, et al.
Published: (2024)
CurriculumLoc: Enhancing Cross-Domain Geolocalization through Multi-Stage Refinement
by: Hu, Boni, et al.
Published: (2023)
by: Hu, Boni, et al.
Published: (2023)
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
by: Wang, Beichen, et al.
Published: (2024)
by: Wang, Beichen, et al.
Published: (2024)
Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
by: Bai, Yongjie, et al.
Published: (2025)
by: Bai, Yongjie, et al.
Published: (2025)
Any-point Trajectory Modeling for Policy Learning
by: Wen, Chuan, et al.
Published: (2023)
by: Wen, Chuan, et al.
Published: (2023)
TeFlow: Enabling Multi-frame Supervision for Self-Supervised Feed-forward Scene Flow Estimation
by: Zhang, Qingwen, et al.
Published: (2026)
by: Zhang, Qingwen, et al.
Published: (2026)
From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving
by: Ang, Sining, et al.
Published: (2026)
by: Ang, Sining, et al.
Published: (2026)
SeePerSea: Multi-modal Perception Dataset of In-water Objects for Autonomous Surface Vehicles
by: Jeong, Mingi, et al.
Published: (2024)
by: Jeong, Mingi, et al.
Published: (2024)
Bridge Thinking and Acting: Unleashing Physical Potential of VLM with Generalizable Action Expert
by: Liu, Mingyu, et al.
Published: (2025)
by: Liu, Mingyu, et al.
Published: (2025)
Neural Radiance Maps for Extraterrestrial Navigation and Path Planning
by: Dai, Adam, et al.
Published: (2026)
by: Dai, Adam, et al.
Published: (2026)
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
by: Shao, Rui, et al.
Published: (2025)
by: Shao, Rui, et al.
Published: (2025)
AdaWorld: Learning Adaptable World Models with Latent Actions
by: Gao, Shenyuan, et al.
Published: (2025)
by: Gao, Shenyuan, et al.
Published: (2025)
See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation
by: Hu, Chih Yao, et al.
Published: (2025)
by: Hu, Chih Yao, et al.
Published: (2025)
KnowVal: A Knowledge-Augmented and Value-Guided Autonomous Driving System
by: Xia, Zhongyu, et al.
Published: (2025)
by: Xia, Zhongyu, et al.
Published: (2025)
BeautyMap: Binary-Encoded Adaptable Ground Matrix for Dynamic Points Removal in Global Maps
by: Jia, Mingkai, et al.
Published: (2024)
by: Jia, Mingkai, et al.
Published: (2024)
Depth Jitter: Seeing through the Depth
by: Rahman, Md Sazidur, et al.
Published: (2025)
by: Rahman, Md Sazidur, et al.
Published: (2025)
Active Next-Best-View Optimization for Risk-Averse Path Planning
by: Khass, Amirhossein Mollaei, et al.
Published: (2025)
by: Khass, Amirhossein Mollaei, et al.
Published: (2025)
DUFOMap: Efficient Dynamic Awareness Mapping
by: Duberg, Daniel, et al.
Published: (2024)
by: Duberg, Daniel, et al.
Published: (2024)
PIGEON: VLM-Driven Object Navigation via Points of Interest Selection
by: Peng, Cheng, et al.
Published: (2025)
by: Peng, Cheng, et al.
Published: (2025)
WorldVLM: Combining World Model Forecasting and Vision-Language Reasoning
by: Englmeier, Stefan, et al.
Published: (2026)
by: Englmeier, Stefan, et al.
Published: (2026)
A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLM
by: Han, ByungOk, et al.
Published: (2024)
by: Han, ByungOk, et al.
Published: (2024)
DNAct: Diffusion Guided Multi-Task 3D Policy Learning
by: Yan, Ge, et al.
Published: (2024)
by: Yan, Ge, et al.
Published: (2024)
Adapt2Reward: Adapting Video-Language Models to Generalizable Robotic Rewards via Failure Prompts
by: Yang, Yanting, et al.
Published: (2024)
by: Yang, Yanting, et al.
Published: (2024)
Similar Items
-
Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
by: Xu, Kechun, et al.
Published: (2025) -
StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
by: Liu, Mingyu, et al.
Published: (2025) -
VLM-Guided Visual Place Recognition for Planet-Scale Geo-Localization
by: Waheed, Sania, et al.
Published: (2025) -
VDT-Auto: End-to-end Autonomous Driving with VLM-Guided Diffusion Transformers
by: Guo, Ziang, et al.
Published: (2025) -
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)