FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Shuang, Chang, Xinyuan, Xie, Mengwei, Liu, Xinran, Bai, Yifan, Pan, Zheng, Xu, Mu, Wei, Xing, Guo, Ning
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918195211796480
author Zeng, Shuang
Chang, Xinyuan
Xie, Mengwei
Liu, Xinran
Bai, Yifan
Pan, Zheng
Xu, Mu
Wei, Xing
Guo, Ning
author_facet Zeng, Shuang
Chang, Xinyuan
Xie, Mengwei
Liu, Xinran
Bai, Yifan
Pan, Zheng
Xu, Mu
Wei, Xing
Guo, Ning
contents Vision-Language-Action (VLA) models offer significant potential for end-to-end driving, yet their reasoning is often constrained by textual Chains-of-Thought (CoT). This symbolic compression of visual information creates a modality gap between perception and planning by blurring spatio-temporal relations and discarding fine-grained cues. We introduce FSDrive, a framework that empowers VLAs to "think visually" using a novel visual spatio-temporal CoT. FSDrive first operates as a world model, generating a unified future frame that combines a predicted background with explicit, physically-plausible priors like future lane dividers and 3D object boxes. This imagined scene serves as the visual spatio-temporal CoT, capturing both spatial structure and temporal evolution in a single representation. The same VLA then functions as an inverse-dynamics model to plan trajectories conditioned on current observations and this visual CoT. We enable this with a unified pre-training paradigm that expands the model's vocabulary with visual tokens and jointly optimizes for semantic understanding (VQA) and future-frame prediction. A progressive curriculum first generates structural priors to enforce physical laws before rendering the full scene. Evaluations on nuScenes and NAVSIM show FSDrive improves trajectory accuracy and reduces collisions, while also achieving competitive FID for video generation with a lightweight autoregressive model and advancing scene understanding on DriveLM. These results confirm that our visual spatio-temporal CoT bridges the perception-planning gap, enabling safer, more anticipatory autonomous driving. Code is available at https://github.com/MIV-XJTU/FSDrive.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17685
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
Zeng, Shuang
Chang, Xinyuan
Xie, Mengwei
Liu, Xinran
Bai, Yifan
Pan, Zheng
Xu, Mu
Wei, Xing
Guo, Ning
Computer Vision and Pattern Recognition
Vision-Language-Action (VLA) models offer significant potential for end-to-end driving, yet their reasoning is often constrained by textual Chains-of-Thought (CoT). This symbolic compression of visual information creates a modality gap between perception and planning by blurring spatio-temporal relations and discarding fine-grained cues. We introduce FSDrive, a framework that empowers VLAs to "think visually" using a novel visual spatio-temporal CoT. FSDrive first operates as a world model, generating a unified future frame that combines a predicted background with explicit, physically-plausible priors like future lane dividers and 3D object boxes. This imagined scene serves as the visual spatio-temporal CoT, capturing both spatial structure and temporal evolution in a single representation. The same VLA then functions as an inverse-dynamics model to plan trajectories conditioned on current observations and this visual CoT. We enable this with a unified pre-training paradigm that expands the model's vocabulary with visual tokens and jointly optimizes for semantic understanding (VQA) and future-frame prediction. A progressive curriculum first generates structural priors to enforce physical laws before rendering the full scene. Evaluations on nuScenes and NAVSIM show FSDrive improves trajectory accuracy and reduces collisions, while also achieving competitive FID for video generation with a lightweight autoregressive model and advancing scene understanding on DriveLM. These results confirm that our visual spatio-temporal CoT bridges the perception-planning gap, enabling safer, more anticipatory autonomous driving. Code is available at https://github.com/MIV-XJTU/FSDrive.
title FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.17685