TAO-Amodal: A Benchmark for Tracking Any Object Amodally

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hsieh, Cheng-Yen, Chen, Kaihua, Dave, Achal, Khurana, Tarasha, Ramanan, Deva
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929301132148736
author Hsieh, Cheng-Yen
Chen, Kaihua
Dave, Achal
Khurana, Tarasha
Ramanan, Deva
author_facet Hsieh, Cheng-Yen
Chen, Kaihua
Dave, Achal
Khurana, Tarasha
Ramanan, Deva
contents Amodal perception, the ability to comprehend complete object structures from partial visibility, is a fundamental skill, even for infants. Its significance extends to applications like autonomous driving, where a clear understanding of heavily occluded objects is essential. However, modern detection and tracking algorithms often overlook this critical capability, perhaps due to the prevalence of \textit{modal} annotations in most benchmarks. To address the scarcity of amodal benchmarks, we introduce TAO-Amodal, featuring 833 diverse categories in thousands of video sequences. Our dataset includes \textit{amodal} and modal bounding boxes for visible and partially or fully occluded objects, including those that are partially out of the camera frame. We investigate the current lay of the land in both amodal tracking and detection by benchmarking state-of-the-art modal trackers and amodal segmentation methods. We find that existing methods, even when adapted for amodal tracking, struggle to detect and track objects under heavy occlusion. To mitigate this, we explore simple finetuning schemes that can increase the amodal tracking and detection metrics of occluded objects by 2.1\% and 3.3\%.
format Preprint
id arxiv_https___arxiv_org_abs_2312_12433
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle TAO-Amodal: A Benchmark for Tracking Any Object Amodally
Hsieh, Cheng-Yen
Chen, Kaihua
Dave, Achal
Khurana, Tarasha
Ramanan, Deva
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Amodal perception, the ability to comprehend complete object structures from partial visibility, is a fundamental skill, even for infants. Its significance extends to applications like autonomous driving, where a clear understanding of heavily occluded objects is essential. However, modern detection and tracking algorithms often overlook this critical capability, perhaps due to the prevalence of \textit{modal} annotations in most benchmarks. To address the scarcity of amodal benchmarks, we introduce TAO-Amodal, featuring 833 diverse categories in thousands of video sequences. Our dataset includes \textit{amodal} and modal bounding boxes for visible and partially or fully occluded objects, including those that are partially out of the camera frame. We investigate the current lay of the land in both amodal tracking and detection by benchmarking state-of-the-art modal trackers and amodal segmentation methods. We find that existing methods, even when adapted for amodal tracking, struggle to detect and track objects under heavy occlusion. To mitigate this, we explore simple finetuning schemes that can increase the amodal tracking and detection metrics of occluded objects by 2.1\% and 3.3\%.
title TAO-Amodal: A Benchmark for Tracking Any Object Amodally
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2312.12433