VideoSAM: A Large Vision Foundation Model for High-Speed Video Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Maduabuchi, Chika, Jossou, Ericmoore, Bucci, Matteo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916599707992064
author Maduabuchi, Chika
Jossou, Ericmoore
Bucci, Matteo
author_facet Maduabuchi, Chika
Jossou, Ericmoore
Bucci, Matteo
contents High-speed video (HSV) segmentation is essential for analyzing dynamic physical processes in scientific and industrial applications, such as boiling heat transfer. Existing models like U-Net struggle with generalization and accurately segmenting complex bubble formations. We present VideoSAM, a specialized adaptation of the Segment Anything Model (SAM), fine-tuned on a diverse HSV dataset for phase detection. Through diverse experiments, VideoSAM demonstrates superior performance across four fluid environments -- Water, FC-72, Nitrogen, and Argon -- significantly outperforming U-Net in complex segmentation tasks. In addition to introducing VideoSAM, we contribute an open-source HSV segmentation dataset designed for phase detection, enabling future research in this domain. Our findings underscore VideoSAM's potential to set new standards in robust and accurate HSV segmentation. The code and dataset used in this study are available online at https://github.com/chikap421/videosam.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21304
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VideoSAM: A Large Vision Foundation Model for High-Speed Video Segmentation
Maduabuchi, Chika
Jossou, Ericmoore
Bucci, Matteo
Computer Vision and Pattern Recognition
Machine Learning
High-speed video (HSV) segmentation is essential for analyzing dynamic physical processes in scientific and industrial applications, such as boiling heat transfer. Existing models like U-Net struggle with generalization and accurately segmenting complex bubble formations. We present VideoSAM, a specialized adaptation of the Segment Anything Model (SAM), fine-tuned on a diverse HSV dataset for phase detection. Through diverse experiments, VideoSAM demonstrates superior performance across four fluid environments -- Water, FC-72, Nitrogen, and Argon -- significantly outperforming U-Net in complex segmentation tasks. In addition to introducing VideoSAM, we contribute an open-source HSV segmentation dataset designed for phase detection, enabling future research in this domain. Our findings underscore VideoSAM's potential to set new standards in robust and accurate HSV segmentation. The code and dataset used in this study are available online at https://github.com/chikap421/videosam.
title VideoSAM: A Large Vision Foundation Model for High-Speed Video Segmentation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2410.21304