Self-Supervised Multimodal NeRF for Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sharma, Gaurav, Kothari, Ravi, Schmid, Josef
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911021070811136
author Sharma, Gaurav
Kothari, Ravi
Schmid, Josef
author_facet Sharma, Gaurav
Kothari, Ravi
Schmid, Josef
contents In this paper, we propose a Neural Radiance Fields (NeRF) based framework, referred to as Novel View Synthesis Framework (NVSF). It jointly learns the implicit neural representation of space and time-varying scene for both LiDAR and Camera. We test this on a real-world autonomous driving scenario containing both static and dynamic scenes. Compared to existing multimodal dynamic NeRFs, our framework is self-supervised, thus eliminating the need for 3D labels. For efficient training and faster convergence, we introduce heuristic-based image pixel sampling to focus on pixels with rich information. To preserve the local features of LiDAR points, a Double Gradient based mask is employed. Extensive experiments on the KITTI-360 dataset show that, compared to the baseline models, our framework has reported best performance on both LiDAR and Camera domain. Code of the model is available at https://github.com/gaurav00700/Selfsupervised-NVSF
format Preprint
id arxiv_https___arxiv_org_abs_2506_19615
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Supervised Multimodal NeRF for Autonomous Driving
Sharma, Gaurav
Kothari, Ravi
Schmid, Josef
Computer Vision and Pattern Recognition
In this paper, we propose a Neural Radiance Fields (NeRF) based framework, referred to as Novel View Synthesis Framework (NVSF). It jointly learns the implicit neural representation of space and time-varying scene for both LiDAR and Camera. We test this on a real-world autonomous driving scenario containing both static and dynamic scenes. Compared to existing multimodal dynamic NeRFs, our framework is self-supervised, thus eliminating the need for 3D labels. For efficient training and faster convergence, we introduce heuristic-based image pixel sampling to focus on pixels with rich information. To preserve the local features of LiDAR points, a Double Gradient based mask is employed. Extensive experiments on the KITTI-360 dataset show that, compared to the baseline models, our framework has reported best performance on both LiDAR and Camera domain. Code of the model is available at https://github.com/gaurav00700/Selfsupervised-NVSF
title Self-Supervised Multimodal NeRF for Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.19615