On Moving Object Segmentation from Monocular Video with Transformers

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Homeyer, Christian, Schnörr, Christoph
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915039609356288
author Homeyer, Christian
Schnörr, Christoph
author_facet Homeyer, Christian
Schnörr, Christoph
contents Moving object detection and segmentation from a single moving camera is a challenging task, requiring an understanding of recognition, motion and 3D geometry. Combining both recognition and reconstruction boils down to a fusion problem, where appearance and motion features need to be combined for classification and segmentation. In this paper, we present a novel fusion architecture for monocular motion segmentation - M3Former, which leverages the strong performance of transformers for segmentation and multi-modal fusion. As reconstructing motion from monocular video is ill-posed, we systematically analyze different 2D and 3D motion representations for this problem and their importance for segmentation performance. Finally, we analyze the effect of training data and show that diverse datasets are required to achieve SotA performance on Kitti and Davis.
format Preprint
id arxiv_https___arxiv_org_abs_2411_19141
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On Moving Object Segmentation from Monocular Video with Transformers
Homeyer, Christian
Schnörr, Christoph
Computer Vision and Pattern Recognition
Artificial Intelligence
Moving object detection and segmentation from a single moving camera is a challenging task, requiring an understanding of recognition, motion and 3D geometry. Combining both recognition and reconstruction boils down to a fusion problem, where appearance and motion features need to be combined for classification and segmentation. In this paper, we present a novel fusion architecture for monocular motion segmentation - M3Former, which leverages the strong performance of transformers for segmentation and multi-modal fusion. As reconstructing motion from monocular video is ill-posed, we systematically analyze different 2D and 3D motion representations for this problem and their importance for segmentation performance. Finally, we analyze the effect of training data and show that diverse datasets are required to achieve SotA performance on Kitti and Davis.
title On Moving Object Segmentation from Monocular Video with Transformers
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.19141