SOFI: Multi-Scale Deformable Transformer for Camera Calibration with Enhanced Line Queries

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Janampa, Sebastian, Pattichis, Marios
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916408569364480
author Janampa, Sebastian
Pattichis, Marios
author_facet Janampa, Sebastian
Pattichis, Marios
contents Camera calibration consists of estimating camera parameters such as the zenith vanishing point and horizon line. Estimating the camera parameters allows other tasks like 3D rendering, artificial reality effects, and object insertion in an image. Transformer-based models have provided promising results; however, they lack cross-scale interaction. In this work, we introduce \textit{multi-Scale defOrmable transFormer for camera calibratIon with enhanced line queries}, SOFI. SOFI improves the line queries used in CTRL-C and MSCC by using both line content and line geometric features. Moreover, SOFI's line queries allow transformer models to adopt the multi-scale deformable attention mechanism to promote cross-scale interaction between the feature maps produced by the backbone. SOFI outperforms existing methods on the \textit {Google Street View}, \textit {Horizon Line in the Wild}, and \textit {Holicity} datasets while keeping a competitive inference speed.
format Preprint
id arxiv_https___arxiv_org_abs_2409_15553
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SOFI: Multi-Scale Deformable Transformer for Camera Calibration with Enhanced Line Queries
Janampa, Sebastian
Pattichis, Marios
Computer Vision and Pattern Recognition
Camera calibration consists of estimating camera parameters such as the zenith vanishing point and horizon line. Estimating the camera parameters allows other tasks like 3D rendering, artificial reality effects, and object insertion in an image. Transformer-based models have provided promising results; however, they lack cross-scale interaction. In this work, we introduce \textit{multi-Scale defOrmable transFormer for camera calibratIon with enhanced line queries}, SOFI. SOFI improves the line queries used in CTRL-C and MSCC by using both line content and line geometric features. Moreover, SOFI's line queries allow transformer models to adopt the multi-scale deformable attention mechanism to promote cross-scale interaction between the feature maps produced by the backbone. SOFI outperforms existing methods on the \textit {Google Street View}, \textit {Horizon Line in the Wild}, and \textit {Holicity} datasets while keeping a competitive inference speed.
title SOFI: Multi-Scale Deformable Transformer for Camera Calibration with Enhanced Line Queries
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.15553