S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Yichen, Xu, Runsheng, He, Tong, Hwang, Jyh-Jing, Luo, Katie, Ji, Jingwei, Lin, Hubert, Chen, Letian, Lu, Yiren, Leng, Zhaoqi, Anguelov, Dragomir, Tan, Mingxing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910983864188928
author Xie, Yichen
Xu, Runsheng
He, Tong
Hwang, Jyh-Jing
Luo, Katie
Ji, Jingwei
Lin, Hubert
Chen, Letian
Lu, Yiren
Leng, Zhaoqi
Anguelov, Dragomir
Tan, Mingxing
author_facet Xie, Yichen
Xu, Runsheng
He, Tong
Hwang, Jyh-Jing
Luo, Katie
Ji, Jingwei
Lin, Hubert
Chen, Letian
Lu, Yiren
Leng, Zhaoqi
Anguelov, Dragomir
Tan, Mingxing
contents The latest advancements in multi-modal large language models (MLLMs) have spurred a strong renewed interest in end-to-end motion planning approaches for autonomous driving. Many end-to-end approaches rely on human annotations to learn intermediate perception and prediction tasks, while purely self-supervised approaches--which directly learn from sensor inputs to generate planning trajectories without human annotations often underperform the state of the art. We observe a key gap in the input representation space: end-to-end approaches built on MLLMs are often pretrained with reasoning tasks in 2D image space rather than the native 3D space in which autonomous vehicles plan. To this end, we propose S4-Driver, a scalable self-supervised motion planning algorithm with spatio-temporal visual representation, based on the popular PaLI multimodal large language model. S4-Driver uses a novel sparse volume strategy to seamlessly transform the strong visual representation of MLLMs from perspective view to 3D space without the need to finetune the vision encoder. This representation aggregates multi-view and multi-frame visual inputs and enables better prediction of planning trajectories in 3D space. To validate our method, we run experiments on both nuScenes and Waymo Open Motion Dataset (with in-house camera data). Results show that S4-Driver performs favorably against existing supervised multi-task approaches while requiring no human annotations. It also demonstrates great scalability when pretrained on large volumes of unannotated driving logs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24139
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation
Xie, Yichen
Xu, Runsheng
He, Tong
Hwang, Jyh-Jing
Luo, Katie
Ji, Jingwei
Lin, Hubert
Chen, Letian
Lu, Yiren
Leng, Zhaoqi
Anguelov, Dragomir
Tan, Mingxing
Computer Vision and Pattern Recognition
Artificial Intelligence
The latest advancements in multi-modal large language models (MLLMs) have spurred a strong renewed interest in end-to-end motion planning approaches for autonomous driving. Many end-to-end approaches rely on human annotations to learn intermediate perception and prediction tasks, while purely self-supervised approaches--which directly learn from sensor inputs to generate planning trajectories without human annotations often underperform the state of the art. We observe a key gap in the input representation space: end-to-end approaches built on MLLMs are often pretrained with reasoning tasks in 2D image space rather than the native 3D space in which autonomous vehicles plan. To this end, we propose S4-Driver, a scalable self-supervised motion planning algorithm with spatio-temporal visual representation, based on the popular PaLI multimodal large language model. S4-Driver uses a novel sparse volume strategy to seamlessly transform the strong visual representation of MLLMs from perspective view to 3D space without the need to finetune the vision encoder. This representation aggregates multi-view and multi-frame visual inputs and enables better prediction of planning trajectories in 3D space. To validate our method, we run experiments on both nuScenes and Waymo Open Motion Dataset (with in-house camera data). Results show that S4-Driver performs favorably against existing supervised multi-task approaches while requiring no human annotations. It also demonstrates great scalability when pretrained on large volumes of unannotated driving logs.
title S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.24139