Segment Anything in Light Fields for Real-Time Applications via Constrained Prompting

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Goncharov, Nikolai, Dansereau, Donald G.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913582496612352
author Goncharov, Nikolai
Dansereau, Donald G.
author_facet Goncharov, Nikolai
Dansereau, Donald G.
contents Segmented light field images can serve as a powerful representation in many of computer vision tasks exploiting geometry and appearance of objects, such as object pose tracking. In the light field domain, segmentation presents an additional objective of recognizing the same segment through all the views. Segment Anything Model 2 (SAM 2) allows producing semantically meaningful segments for monocular images and videos. However, using SAM 2 directly on light fields is highly ineffective due to unexploited constraints. In this work, we present a novel light field segmentation method that adapts SAM 2 to the light field domain without retraining or modifying the model. By utilizing the light field domain constraints, the method produces high quality and view-consistent light field masks, outperforming the SAM 2 video tracking baseline and working 7 times faster, with a real-time speed. We achieve this by exploiting the epipolar geometry cues to propagate the masks between the views, probing the SAM 2 latent space to estimate their occlusion, and further prompting SAM 2 for their refinement.
format Preprint
id arxiv_https___arxiv_org_abs_2411_13840
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Segment Anything in Light Fields for Real-Time Applications via Constrained Prompting
Goncharov, Nikolai
Dansereau, Donald G.
Computer Vision and Pattern Recognition
Segmented light field images can serve as a powerful representation in many of computer vision tasks exploiting geometry and appearance of objects, such as object pose tracking. In the light field domain, segmentation presents an additional objective of recognizing the same segment through all the views. Segment Anything Model 2 (SAM 2) allows producing semantically meaningful segments for monocular images and videos. However, using SAM 2 directly on light fields is highly ineffective due to unexploited constraints. In this work, we present a novel light field segmentation method that adapts SAM 2 to the light field domain without retraining or modifying the model. By utilizing the light field domain constraints, the method produces high quality and view-consistent light field masks, outperforming the SAM 2 video tracking baseline and working 7 times faster, with a real-time speed. We achieve this by exploiting the epipolar geometry cues to propagate the masks between the views, probing the SAM 2 latent space to estimate their occlusion, and further prompting SAM 2 for their refinement.
title Segment Anything in Light Fields for Real-Time Applications via Constrained Prompting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.13840