MediViSTA: Medical Video Segmentation via Temporal Fusion SAM Adaptation for Echocardiography

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Sekeun, Jin, Pengfei, Chen, Cheng, Kim, Kyungsang, Lyu, Zhiliang, Ren, Hui, Kim, Sunghwan, Liu, Zhengliang, Zhong, Aoxiao, Liu, Tianming, Li, Xiang, Li, Quanzheng
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913572921016320
author Kim, Sekeun
Jin, Pengfei
Chen, Cheng
Kim, Kyungsang
Lyu, Zhiliang
Ren, Hui
Kim, Sunghwan
Liu, Zhengliang
Zhong, Aoxiao
Liu, Tianming
Li, Xiang
Li, Quanzheng
author_facet Kim, Sekeun
Jin, Pengfei
Chen, Cheng
Kim, Kyungsang
Lyu, Zhiliang
Ren, Hui
Kim, Sunghwan
Liu, Zhengliang
Zhong, Aoxiao
Liu, Tianming
Li, Xiang
Li, Quanzheng
contents Despite achieving impressive results in general-purpose semantic segmentation with strong generalization on natural images, the Segment Anything Model (SAM) has shown less precision and stability in medical image segmentation. In particular, the original SAM architecture is designed for 2D natural images and is therefore not support to handle three-dimensional information, which is particularly important for medical imaging modalities that are often volumetric or video data. In this paper, we introduce MediViSTA, a parameter-efficient fine-tuning method designed to adapt the vision foundation model for medical video, with a specific focus on echocardiographic segmentation. To achieve spatial adaptation, we propose a frequency feature fusion technique that injects spatial frequency information from a CNN branch. For temporal adaptation, we integrate temporal adapters within the transformer blocks of the image encoder. Using a fine-tuning strategy, only a small subset of pre-trained parameters is updated, allowing efficient adaptation to echocardiographic data. The effectiveness of our method has been comprehensively evaluated on three datasets, comprising two public datasets and one multi-center in-house dataset. Our method consistently outperforms various state-of-the-art approaches without using any prompts. Furthermore, our model exhibits strong generalization capabilities on unseen datasets, surpassing the second-best approach by 2.15\% in Dice and 0.09 in temporal consistency. The results demonstrate the potential of MediViSTA to significantly advance echocardiographical video segmentation, offering improved accuracy and robustness in cardiac assessment applications.
format Preprint
id arxiv_https___arxiv_org_abs_2309_13539
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MediViSTA: Medical Video Segmentation via Temporal Fusion SAM Adaptation for Echocardiography
Kim, Sekeun
Jin, Pengfei
Chen, Cheng
Kim, Kyungsang
Lyu, Zhiliang
Ren, Hui
Kim, Sunghwan
Liu, Zhengliang
Zhong, Aoxiao
Liu, Tianming
Li, Xiang
Li, Quanzheng
Image and Video Processing
Despite achieving impressive results in general-purpose semantic segmentation with strong generalization on natural images, the Segment Anything Model (SAM) has shown less precision and stability in medical image segmentation. In particular, the original SAM architecture is designed for 2D natural images and is therefore not support to handle three-dimensional information, which is particularly important for medical imaging modalities that are often volumetric or video data. In this paper, we introduce MediViSTA, a parameter-efficient fine-tuning method designed to adapt the vision foundation model for medical video, with a specific focus on echocardiographic segmentation. To achieve spatial adaptation, we propose a frequency feature fusion technique that injects spatial frequency information from a CNN branch. For temporal adaptation, we integrate temporal adapters within the transformer blocks of the image encoder. Using a fine-tuning strategy, only a small subset of pre-trained parameters is updated, allowing efficient adaptation to echocardiographic data. The effectiveness of our method has been comprehensively evaluated on three datasets, comprising two public datasets and one multi-center in-house dataset. Our method consistently outperforms various state-of-the-art approaches without using any prompts. Furthermore, our model exhibits strong generalization capabilities on unseen datasets, surpassing the second-best approach by 2.15\% in Dice and 0.09 in temporal consistency. The results demonstrate the potential of MediViSTA to significantly advance echocardiographical video segmentation, offering improved accuracy and robustness in cardiac assessment applications.
title MediViSTA: Medical Video Segmentation via Temporal Fusion SAM Adaptation for Echocardiography
topic Image and Video Processing
url https://arxiv.org/abs/2309.13539