Salvato in:
Dettagli Bibliografici
Autori principali: Hudson, Finlay G. C., Smith, William A. P.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:https://arxiv.org/abs/2411.19210
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911487050645504
author Hudson, Finlay G. C.
Smith, William A. P.
author_facet Hudson, Finlay G. C.
Smith, William A. P.
contents We present Track Anything Behind Everything (TABE), a novel pipeline for zero-shot amodal video object segmentation. Unlike existing methods that require pretrained class labels, our approach uses a single query mask from the first frame where the object is visible, enabling flexible, zero-shot inference. We pose amodal segmentation as generative outpainting from modal (visible) masks using a pretrained video diffusion model. We do not need to re-train the diffusion model to accommodate additional input channels but instead use a pretrained model that we fine-tune at test-time to allow specialisation towards the tracked object. Our TABE pipeline is specifically designed to handle amodal completion, even in scenarios where objects are completely occluded. Our model and code will all be released.
format Preprint
id arxiv_https___arxiv_org_abs_2411_19210
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Track Anything Behind Everything: Zero-Shot Amodal Video Object Segmentation
Hudson, Finlay G. C.
Smith, William A. P.
Computer Vision and Pattern Recognition
We present Track Anything Behind Everything (TABE), a novel pipeline for zero-shot amodal video object segmentation. Unlike existing methods that require pretrained class labels, our approach uses a single query mask from the first frame where the object is visible, enabling flexible, zero-shot inference. We pose amodal segmentation as generative outpainting from modal (visible) masks using a pretrained video diffusion model. We do not need to re-train the diffusion model to accommodate additional input channels but instead use a pretrained model that we fine-tune at test-time to allow specialisation towards the tracked object. Our TABE pipeline is specifically designed to handle amodal completion, even in scenarios where objects are completely occluded. Our model and code will all be released.
title Track Anything Behind Everything: Zero-Shot Amodal Video Object Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.19210