GASP: Unifying Geometric and Semantic Self-Supervised Pre-training for Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ljungbergh, William, Lilja, Adam, Ling, Adam Tonderski. Arvid Laveno, Lindström, Carl, Verbeke, Willem, Fu, Junsheng, Petersson, Christoffer, Hammarstrand, Lars, Felsberg, Michael
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912283907588096
author Ljungbergh, William
Lilja, Adam
Ling, Adam Tonderski. Arvid Laveno
Lindström, Carl
Verbeke, Willem
Fu, Junsheng
Petersson, Christoffer
Hammarstrand, Lars
Felsberg, Michael
author_facet Ljungbergh, William
Lilja, Adam
Ling, Adam Tonderski. Arvid Laveno
Lindström, Carl
Verbeke, Willem
Fu, Junsheng
Petersson, Christoffer
Hammarstrand, Lars
Felsberg, Michael
contents Self-supervised pre-training based on next-token prediction has enabled large language models to capture the underlying structure of text, and has led to unprecedented performance on a large array of tasks when applied at scale. Similarly, autonomous driving generates vast amounts of spatiotemporal data, alluding to the possibility of harnessing scale to learn the underlying geometric and semantic structure of the environment and its evolution over time. In this direction, we propose a geometric and semantic self-supervised pre-training method, GASP, that learns a unified representation by predicting, at any queried future point in spacetime, (1) general occupancy, capturing the evolving structure of the 3D scene; (2) ego occupancy, modeling the ego vehicle path through the environment; and (3) distilled high-level features from a vision foundation model. By modeling geometric and semantic 4D occupancy fields instead of raw sensor measurements, the model learns a structured, generalizable representation of the environment and its evolution through time. We validate GASP on multiple autonomous driving benchmarks, demonstrating significant improvements in semantic occupancy forecasting, online mapping, and ego trajectory prediction. Our results demonstrate that continuous 4D geometric and semantic occupancy prediction provides a scalable and effective pre-training paradigm for autonomous driving. For code and additional visualizations, see \href{https://research.zenseact.com/publications/gasp/.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15672
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GASP: Unifying Geometric and Semantic Self-Supervised Pre-training for Autonomous Driving
Ljungbergh, William
Lilja, Adam
Ling, Adam Tonderski. Arvid Laveno
Lindström, Carl
Verbeke, Willem
Fu, Junsheng
Petersson, Christoffer
Hammarstrand, Lars
Felsberg, Michael
Computer Vision and Pattern Recognition
Robotics
Self-supervised pre-training based on next-token prediction has enabled large language models to capture the underlying structure of text, and has led to unprecedented performance on a large array of tasks when applied at scale. Similarly, autonomous driving generates vast amounts of spatiotemporal data, alluding to the possibility of harnessing scale to learn the underlying geometric and semantic structure of the environment and its evolution over time. In this direction, we propose a geometric and semantic self-supervised pre-training method, GASP, that learns a unified representation by predicting, at any queried future point in spacetime, (1) general occupancy, capturing the evolving structure of the 3D scene; (2) ego occupancy, modeling the ego vehicle path through the environment; and (3) distilled high-level features from a vision foundation model. By modeling geometric and semantic 4D occupancy fields instead of raw sensor measurements, the model learns a structured, generalizable representation of the environment and its evolution through time. We validate GASP on multiple autonomous driving benchmarks, demonstrating significant improvements in semantic occupancy forecasting, online mapping, and ego trajectory prediction. Our results demonstrate that continuous 4D geometric and semantic occupancy prediction provides a scalable and effective pre-training paradigm for autonomous driving. For code and additional visualizations, see \href{https://research.zenseact.com/publications/gasp/.
title GASP: Unifying Geometric and Semantic Self-Supervised Pre-training for Autonomous Driving
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2503.15672