STELLAR: Scaling 3D Perception Large Models for Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yingwei, Huang, Xin, Liu, Yang, Fu, Yang, Zhu, Alex Zihao, Song, Chen, Yao, Junwen, Subramanian, Anant, Xiang, Hao, Shi, Weijing, Zou, Yuliang, Hoddes, Tom, Leng, Zhaoqi, Thattai, Govind, Anguelov, Dragomir, Tan, Mingxing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914581429878784
author Li, Yingwei
Huang, Xin
Liu, Yang
Fu, Yang
Zhu, Alex Zihao
Song, Chen
Yao, Junwen
Subramanian, Anant
Xiang, Hao
Shi, Weijing
Zou, Yuliang
Hoddes, Tom
Leng, Zhaoqi
Thattai, Govind
Anguelov, Dragomir
Tan, Mingxing
author_facet Li, Yingwei
Huang, Xin
Liu, Yang
Fu, Yang
Zhu, Alex Zihao
Song, Chen
Yao, Junwen
Subramanian, Anant
Xiang, Hao
Shi, Weijing
Zou, Yuliang
Hoddes, Tom
Leng, Zhaoqi
Thattai, Govind
Anguelov, Dragomir
Tan, Mingxing
contents Model scaling has demonstrated remarkable success through large-scale training on diverse datasets. It remains an open question whether the same paradigm would apply to autonomous driving perception systems due to unique challenges, such as fusing heterogeneous sensor data and the need for sophisticated 3D spatial understanding. To bridge this gap, we present a comprehensive study on systematically analyzing the impact of scale on these systems. We develop our STELLAR model based on Sparse Window Transformer, by extending the input modalities to include LiDAR, radar, camera, and map prior. We train the model on a large-scale dataset of 50 million driving examples with up to 500 million parameters. Our large-scale experiments reveal empirical scaling trends that connect model performance to model size, data, and compute. The resulting model establishes a new state-of-the-art on the Waymo Open Dataset challenge, outperforming prior arts by a large margin. Our work demonstrates that large-scale training is a highly promising path for advancing the capabilities of perception models for autonomous driving.
format Preprint
id arxiv_https___arxiv_org_abs_2605_20390
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle STELLAR: Scaling 3D Perception Large Models for Autonomous Driving
Li, Yingwei
Huang, Xin
Liu, Yang
Fu, Yang
Zhu, Alex Zihao
Song, Chen
Yao, Junwen
Subramanian, Anant
Xiang, Hao
Shi, Weijing
Zou, Yuliang
Hoddes, Tom
Leng, Zhaoqi
Thattai, Govind
Anguelov, Dragomir
Tan, Mingxing
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
Model scaling has demonstrated remarkable success through large-scale training on diverse datasets. It remains an open question whether the same paradigm would apply to autonomous driving perception systems due to unique challenges, such as fusing heterogeneous sensor data and the need for sophisticated 3D spatial understanding. To bridge this gap, we present a comprehensive study on systematically analyzing the impact of scale on these systems. We develop our STELLAR model based on Sparse Window Transformer, by extending the input modalities to include LiDAR, radar, camera, and map prior. We train the model on a large-scale dataset of 50 million driving examples with up to 500 million parameters. Our large-scale experiments reveal empirical scaling trends that connect model performance to model size, data, and compute. The resulting model establishes a new state-of-the-art on the Waymo Open Dataset challenge, outperforming prior arts by a large margin. Our work demonstrates that large-scale training is a highly promising path for advancing the capabilities of perception models for autonomous driving.
title STELLAR: Scaling 3D Perception Large Models for Autonomous Driving
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2605.20390