SMITE: Segment Me In TimE

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alimohammadi, Amirhossein, Nag, Sauradip, Taghanaki, Saeid Asgari, Tagliasacchi, Andrea, Hamarneh, Ghassan, Amiri, Ali Mahdavi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913697877721088
author Alimohammadi, Amirhossein
Nag, Sauradip
Taghanaki, Saeid Asgari
Tagliasacchi, Andrea
Hamarneh, Ghassan
Amiri, Ali Mahdavi
author_facet Alimohammadi, Amirhossein
Nag, Sauradip
Taghanaki, Saeid Asgari
Tagliasacchi, Andrea
Hamarneh, Ghassan
Amiri, Ali Mahdavi
contents Segmenting an object in a video presents significant challenges. Each pixel must be accurately labelled, and these labels must remain consistent across frames. The difficulty increases when the segmentation is with arbitrary granularity, meaning the number of segments can vary arbitrarily, and masks are defined based on only one or a few sample images. In this paper, we address this issue by employing a pre-trained text to image diffusion model supplemented with an additional tracking mechanism. We demonstrate that our approach can effectively manage various segmentation scenarios and outperforms state-of-the-art alternatives.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18538
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SMITE: Segment Me In TimE
Alimohammadi, Amirhossein
Nag, Sauradip
Taghanaki, Saeid Asgari
Tagliasacchi, Andrea
Hamarneh, Ghassan
Amiri, Ali Mahdavi
Computer Vision and Pattern Recognition
Segmenting an object in a video presents significant challenges. Each pixel must be accurately labelled, and these labels must remain consistent across frames. The difficulty increases when the segmentation is with arbitrary granularity, meaning the number of segments can vary arbitrarily, and masks are defined based on only one or a few sample images. In this paper, we address this issue by employing a pre-trained text to image diffusion model supplemented with an additional tracking mechanism. We demonstrate that our approach can effectively manage various segmentation scenarios and outperforms state-of-the-art alternatives.
title SMITE: Segment Me In TimE
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.18538