Robustness Analysis on Foundational Segmentation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schiappa, Madeline Chantry, Azad, Shehreen, VS, Sachidanand, Ge, Yunhao, Miksik, Ondrej, Rawat, Yogesh S., Vineet, Vibhav
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913332040040448
author Schiappa, Madeline Chantry
Azad, Shehreen
VS, Sachidanand
Ge, Yunhao
Miksik, Ondrej
Rawat, Yogesh S.
Vineet, Vibhav
author_facet Schiappa, Madeline Chantry
Azad, Shehreen
VS, Sachidanand
Ge, Yunhao
Miksik, Ondrej
Rawat, Yogesh S.
Vineet, Vibhav
contents Due to the increase in computational resources and accessibility of data, an increase in large, deep learning models trained on copious amounts of multi-modal data using self-supervised or semi-supervised learning have emerged. These ``foundation'' models are often adapted to a variety of downstream tasks like classification, object detection, and segmentation with little-to-no training on the target dataset. In this work, we perform a robustness analysis of Visual Foundation Models (VFMs) for segmentation tasks and focus on robustness against real-world distribution shift inspired perturbations. We benchmark seven state-of-the-art segmentation architectures using 2 different perturbed datasets, MS COCO-P and ADE20K-P, with 17 different perturbations with 5 severity levels each. Our findings reveal several key insights: (1) VFMs exhibit vulnerabilities to compression-induced corruptions, (2) despite not outpacing all of unimodal models in robustness, multimodal models show competitive resilience in zero-shot scenarios, and (3) VFMs demonstrate enhanced robustness for certain object categories. These observations suggest that our robustness evaluation framework sets new requirements for foundational models, encouraging further advancements to bolster their adaptability and performance. The code and dataset is available at: \url{https://tinyurl.com/fm-robust}.
format Preprint
id arxiv_https___arxiv_org_abs_2306_09278
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Robustness Analysis on Foundational Segmentation Models
Schiappa, Madeline Chantry
Azad, Shehreen
VS, Sachidanand
Ge, Yunhao
Miksik, Ondrej
Rawat, Yogesh S.
Vineet, Vibhav
Computer Vision and Pattern Recognition
Due to the increase in computational resources and accessibility of data, an increase in large, deep learning models trained on copious amounts of multi-modal data using self-supervised or semi-supervised learning have emerged. These ``foundation'' models are often adapted to a variety of downstream tasks like classification, object detection, and segmentation with little-to-no training on the target dataset. In this work, we perform a robustness analysis of Visual Foundation Models (VFMs) for segmentation tasks and focus on robustness against real-world distribution shift inspired perturbations. We benchmark seven state-of-the-art segmentation architectures using 2 different perturbed datasets, MS COCO-P and ADE20K-P, with 17 different perturbations with 5 severity levels each. Our findings reveal several key insights: (1) VFMs exhibit vulnerabilities to compression-induced corruptions, (2) despite not outpacing all of unimodal models in robustness, multimodal models show competitive resilience in zero-shot scenarios, and (3) VFMs demonstrate enhanced robustness for certain object categories. These observations suggest that our robustness evaluation framework sets new requirements for foundational models, encouraging further advancements to bolster their adaptability and performance. The code and dataset is available at: \url{https://tinyurl.com/fm-robust}.
title Robustness Analysis on Foundational Segmentation Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2306.09278