STELLAR: Scaling 3D Perception Large Models for Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914581429878784 |
|---|---|
| author | Li, Yingwei Huang, Xin Liu, Yang Fu, Yang Zhu, Alex Zihao Song, Chen Yao, Junwen Subramanian, Anant Xiang, Hao Shi, Weijing Zou, Yuliang Hoddes, Tom Leng, Zhaoqi Thattai, Govind Anguelov, Dragomir Tan, Mingxing |
| author_facet | Li, Yingwei Huang, Xin Liu, Yang Fu, Yang Zhu, Alex Zihao Song, Chen Yao, Junwen Subramanian, Anant Xiang, Hao Shi, Weijing Zou, Yuliang Hoddes, Tom Leng, Zhaoqi Thattai, Govind Anguelov, Dragomir Tan, Mingxing |
| contents | Model scaling has demonstrated remarkable success through large-scale training on diverse datasets. It remains an open question whether the same paradigm would apply to autonomous driving perception systems due to unique challenges, such as fusing heterogeneous sensor data and the need for sophisticated 3D spatial understanding. To bridge this gap, we present a comprehensive study on systematically analyzing the impact of scale on these systems. We develop our STELLAR model based on Sparse Window Transformer, by extending the input modalities to include LiDAR, radar, camera, and map prior. We train the model on a large-scale dataset of 50 million driving examples with up to 500 million parameters. Our large-scale experiments reveal empirical scaling trends that connect model performance to model size, data, and compute. The resulting model establishes a new state-of-the-art on the Waymo Open Dataset challenge, outperforming prior arts by a large margin. Our work demonstrates that large-scale training is a highly promising path for advancing the capabilities of perception models for autonomous driving. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_20390 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | STELLAR: Scaling 3D Perception Large Models for Autonomous Driving Li, Yingwei Huang, Xin Liu, Yang Fu, Yang Zhu, Alex Zihao Song, Chen Yao, Junwen Subramanian, Anant Xiang, Hao Shi, Weijing Zou, Yuliang Hoddes, Tom Leng, Zhaoqi Thattai, Govind Anguelov, Dragomir Tan, Mingxing Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Robotics Model scaling has demonstrated remarkable success through large-scale training on diverse datasets. It remains an open question whether the same paradigm would apply to autonomous driving perception systems due to unique challenges, such as fusing heterogeneous sensor data and the need for sophisticated 3D spatial understanding. To bridge this gap, we present a comprehensive study on systematically analyzing the impact of scale on these systems. We develop our STELLAR model based on Sparse Window Transformer, by extending the input modalities to include LiDAR, radar, camera, and map prior. We train the model on a large-scale dataset of 50 million driving examples with up to 500 million parameters. Our large-scale experiments reveal empirical scaling trends that connect model performance to model size, data, and compute. The resulting model establishes a new state-of-the-art on the Waymo Open Dataset challenge, outperforming prior arts by a large margin. Our work demonstrates that large-scale training is a highly promising path for advancing the capabilities of perception models for autonomous driving. |
| title | STELLAR: Scaling 3D Perception Large Models for Autonomous Driving |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Robotics |
| url | https://arxiv.org/abs/2605.20390 |