CalibFormer: A Transformer-based Automatic LiDAR-Camera Calibration Network

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Yuxuan, Li, Yao, Meng, Chengzhen, Li, Xingchen, Ji, Jianmin, Zhang, Yanyong
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909138829705216
author Xiao, Yuxuan
Li, Yao
Meng, Chengzhen
Li, Xingchen
Ji, Jianmin
Zhang, Yanyong
author_facet Xiao, Yuxuan
Li, Yao
Meng, Chengzhen
Li, Xingchen
Ji, Jianmin
Zhang, Yanyong
contents The fusion of LiDARs and cameras has been increasingly adopted in autonomous driving for perception tasks. The performance of such fusion-based algorithms largely depends on the accuracy of sensor calibration, which is challenging due to the difficulty of identifying common features across different data modalities. Previously, many calibration methods involved specific targets and/or manual intervention, which has proven to be cumbersome and costly. Learning-based online calibration methods have been proposed, but their performance is barely satisfactory in most cases. These methods usually suffer from issues such as sparse feature maps, unreliable cross-modality association, inaccurate calibration parameter regression, etc. In this paper, to address these issues, we propose CalibFormer, an end-to-end network for automatic LiDAR-camera calibration. We aggregate multiple layers of camera and LiDAR image features to achieve high-resolution representations. A multi-head correlation module is utilized to identify correlations between features more accurately. Lastly, we employ transformer architectures to estimate accurate calibration parameters from the correlation information. Our method achieved a mean translation error of $0.8751 \mathrm{cm}$ and a mean rotation error of $0.0562 ^{\circ}$ on the KITTI dataset, surpassing existing state-of-the-art methods and demonstrating strong robustness, accuracy, and generalization capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2311_15241
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle CalibFormer: A Transformer-based Automatic LiDAR-Camera Calibration Network
Xiao, Yuxuan
Li, Yao
Meng, Chengzhen
Li, Xingchen
Ji, Jianmin
Zhang, Yanyong
Computer Vision and Pattern Recognition
Robotics
The fusion of LiDARs and cameras has been increasingly adopted in autonomous driving for perception tasks. The performance of such fusion-based algorithms largely depends on the accuracy of sensor calibration, which is challenging due to the difficulty of identifying common features across different data modalities. Previously, many calibration methods involved specific targets and/or manual intervention, which has proven to be cumbersome and costly. Learning-based online calibration methods have been proposed, but their performance is barely satisfactory in most cases. These methods usually suffer from issues such as sparse feature maps, unreliable cross-modality association, inaccurate calibration parameter regression, etc. In this paper, to address these issues, we propose CalibFormer, an end-to-end network for automatic LiDAR-camera calibration. We aggregate multiple layers of camera and LiDAR image features to achieve high-resolution representations. A multi-head correlation module is utilized to identify correlations between features more accurately. Lastly, we employ transformer architectures to estimate accurate calibration parameters from the correlation information. Our method achieved a mean translation error of $0.8751 \mathrm{cm}$ and a mean rotation error of $0.0562 ^{\circ}$ on the KITTI dataset, surpassing existing state-of-the-art methods and demonstrating strong robustness, accuracy, and generalization capabilities.
title CalibFormer: A Transformer-based Automatic LiDAR-Camera Calibration Network
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2311.15241