S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Yichen, Xu, Runsheng, He, Tong, Hwang, Jyh-Jing, Luo, Katie, Ji, Jingwei, Lin, Hubert, Chen, Letian, Lu, Yiren, Leng, Zhaoqi, Anguelov, Dragomir, Tan, Mingxing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
by: Luo, Katie, et al.
Published: (2025)
by: Luo, Katie, et al.
Published: (2025)
PVTransformer: Point-to-Voxel Transformer for Scalable 3D Object Detection
by: Leng, Zhaoqi, et al.
Published: (2024)
by: Leng, Zhaoqi, et al.
Published: (2024)
EMMA: End-to-End Multimodal Model for Autonomous Driving
by: Hwang, Jyh-Jing, et al.
Published: (2024)
by: Hwang, Jyh-Jing, et al.
Published: (2024)
WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios
by: Xu, Runsheng, et al.
Published: (2025)
by: Xu, Runsheng, et al.
Published: (2025)
LET-3D-AP: Longitudinal Error Tolerant 3D Average Precision for Camera-Only 3D Detection
by: Hung, Wei-Chih, et al.
Published: (2022)
by: Hung, Wei-Chih, et al.
Published: (2022)
Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
STELLAR: Scaling 3D Perception Large Models for Autonomous Driving
by: Li, Yingwei, et al.
Published: (2026)
by: Li, Yingwei, et al.
Published: (2026)
SceneCrafter: Controllable Multi-View Driving Scene Editing
by: Zhu, Zehao, et al.
Published: (2025)
by: Zhu, Zehao, et al.
Published: (2025)
SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model
by: Tan, Shuhan, et al.
Published: (2025)
by: Tan, Shuhan, et al.
Published: (2025)
WOMD-LiDAR: Raw Sensor Dataset Benchmark for Motion Forecasting
by: Chen, Kan, et al.
Published: (2023)
by: Chen, Kan, et al.
Published: (2023)
SceneDiffuser: Efficient and Controllable Driving Simulation Initialization and Rollout
by: Jiang, Chiyu Max, et al.
Published: (2024)
by: Jiang, Chiyu Max, et al.
Published: (2024)
Identifying Spatio-Temporal Drivers of Extreme Events
by: Eddin, Mohamad Hakam Shams, et al.
Published: (2024)
by: Eddin, Mohamad Hakam Shams, et al.
Published: (2024)
Temporal Embeddings: Scalable Self-Supervised Temporal Representation Learning from Spatiotemporal Data for Multimodal Computer Vision
by: Cao, Yi, et al.
Published: (2023)
by: Cao, Yi, et al.
Published: (2023)
Scene Reconstruction as Mapping Priors for 3D Detection
by: Fu, Yang, et al.
Published: (2026)
by: Fu, Yang, et al.
Published: (2026)
State‐owned Enterprises and the Politics of Financializing Infrastructure Development in Indonesia: De‐risking at the Limit?
by: Dimitar Anguelov
Published: (2024)
by: Dimitar Anguelov
Published: (2024)
Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving
by: Wang, Jiahao, et al.
Published: (2026)
by: Wang, Jiahao, et al.
Published: (2026)
Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
by: Guo, Xiangyu, et al.
Published: (2025)
by: Guo, Xiangyu, et al.
Published: (2025)
DriveFix: Spatio-Temporally Coherent Driving Scene Restoration
by: Si, Heyu, et al.
Published: (2026)
by: Si, Heyu, et al.
Published: (2026)
Context-Aware Multimodal Representation Learning for Spatio-Temporally Explicit Environmental Modelling
by: Peters, Julia, et al.
Published: (2025)
by: Peters, Julia, et al.
Published: (2025)
Nl2Hltl2Plan: Scaling Up Natural Language Understanding for Multi-Robots Through Hierarchical Temporal Logic Task Representation
by: Xu, Shaojun, et al.
Published: (2024)
by: Xu, Shaojun, et al.
Published: (2024)
Geolocation Representation from Large Language Models are Generic Enhancers for Spatio-Temporal Learning
by: He, Junlin, et al.
Published: (2024)
by: He, Junlin, et al.
Published: (2024)
Coastal Hypoxia in the Indian Ocean: Unraveling Drivers of Spatio‐Temporal Variability
by: Fan Yang, et al.
Published: (2025)
by: Fan Yang, et al.
Published: (2025)
A Spatio-Temporal Representation Learning as an Alternative to Traditional Glosses in Sign Language Translation and Production
by: Hwang, Eui Jun, et al.
Published: (2024)
by: Hwang, Eui Jun, et al.
Published: (2024)
Bayesian Design for Sampling Anomalous Spatio-Temporal Data
by: Buchhorn, Katie, et al.
Published: (2024)
by: Buchhorn, Katie, et al.
Published: (2024)
AdaSTI: Conditional Diffusion Models with Adaptive Dependency Modeling for Spatio-Temporal Imputation
by: Yang, Yubo, et al.
Published: (2025)
by: Yang, Yubo, et al.
Published: (2025)
FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
by: Zeng, Shuang, et al.
Published: (2025)
by: Zeng, Shuang, et al.
Published: (2025)
U-STS-LLM A Unified Spatio-Temporal Steered Large Language Model for Traffic Prediction and Imputation
by: Zhang, Yichen, et al.
Published: (2026)
by: Zhang, Yichen, et al.
Published: (2026)
Simulation Design for Velocity-Controlled Spatio-Temporal Drivers in Laser Wakefield Acceleration
by: Badiali, Chiara, et al.
Published: (2026)
by: Badiali, Chiara, et al.
Published: (2026)
STDR: Spatio-Temporal Decoupling for Real-Time Dynamic Scene Rendering
by: Li, Zehao, et al.
Published: (2025)
by: Li, Zehao, et al.
Published: (2025)
Stationary states of aggregation-diffusion equations with compactly supported attraction kernels: radial symmetry and mass-independent boundedness
by: Anguelov, Roumen, et al.
Published: (2024)
by: Anguelov, Roumen, et al.
Published: (2024)
STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding
by: Garg, Aaryan, et al.
Published: (2025)
by: Garg, Aaryan, et al.
Published: (2025)
Flexible and Scalable Bayesian Modelling of Spatio-Temporal Hawkes Processes
by: Liu, Wenqing, et al.
Published: (2026)
by: Liu, Wenqing, et al.
Published: (2026)
Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking
by: Zheng, Yaozong, et al.
Published: (2025)
by: Zheng, Yaozong, et al.
Published: (2025)
Spatio-Temporal Self-Supervised Learning for Traffic Flow Prediction
by: Ji, Jiahao, et al.
Published: (2022)
by: Ji, Jiahao, et al.
Published: (2022)
Beyond Spatio-Temporal Representations: Evolving Fourier Transform for Temporal Graphs
by: Bastos, Anson, et al.
Published: (2024)
by: Bastos, Anson, et al.
Published: (2024)
A Distributed Hierarchical Spatio-Temporal Edge-Enhanced Graph Neural Network for City-Scale Dynamic Logistics Routing
by: Han, Zihan, et al.
Published: (2025)
by: Han, Zihan, et al.
Published: (2025)
Vectorized Video Representation with Easy Editing via Hierarchical Spatio-Temporally Consistent Proxy Embedding
by: Chen, Ye, et al.
Published: (2025)
by: Chen, Ye, et al.
Published: (2025)
Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders
by: Dokme, Atahan, et al.
Published: (2026)
by: Dokme, Atahan, et al.
Published: (2026)
Transformer RGBT Tracking with Spatio-Temporal Multimodal Tokens
by: Sun, Dengdi, et al.
Published: (2024)
by: Sun, Dengdi, et al.
Published: (2024)
STDA: Spatio-Temporal Dual-Encoder Network Incorporating Driver Attention to Predict Driver Behaviors Under Safety-Critical Scenarios
by: Xu, Dongyang, et al.
Published: (2024)
by: Xu, Dongyang, et al.
Published: (2024)
Similar Items
-
Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
by: Luo, Katie, et al.
Published: (2025) -
PVTransformer: Point-to-Voxel Transformer for Scalable 3D Object Detection
by: Leng, Zhaoqi, et al.
Published: (2024) -
EMMA: End-to-End Multimodal Model for Autonomous Driving
by: Hwang, Jyh-Jing, et al.
Published: (2024) -
WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios
by: Xu, Runsheng, et al.
Published: (2025) -
LET-3D-AP: Longitudinal Error Tolerant 3D Average Precision for Camera-Only 3D Detection
by: Hung, Wei-Chih, et al.
Published: (2022)