FASL-Seg: Anatomy and Tool Segmentation of Surgical Scenes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abdel-Ghani, Muraam, Ali, Mahmoud, Ali, Mohamed, Ahmed, Fatmaelzahraa, Arsalan, Muhammad, Al-Ali, Abdulaziz, Balakrishnan, Shidin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914123193778176
author Abdel-Ghani, Muraam
Ali, Mahmoud
Ali, Mohamed
Ahmed, Fatmaelzahraa
Arsalan, Muhammad
Al-Ali, Abdulaziz
Balakrishnan, Shidin
author_facet Abdel-Ghani, Muraam
Ali, Mahmoud
Ali, Mohamed
Ahmed, Fatmaelzahraa
Arsalan, Muhammad
Al-Ali, Abdulaziz
Balakrishnan, Shidin
contents The growing popularity of robotic minimally invasive surgeries has made deep learning-based surgical training a key area of research. A thorough understanding of the surgical scene components is crucial, which semantic segmentation models can help achieve. However, most existing work focuses on surgical tools and overlooks anatomical objects. Additionally, current state-of-the-art (SOTA) models struggle to balance capturing high-level contextual features and low-level edge features. We propose a Feature-Adaptive Spatial Localization model (FASL-Seg), designed to capture features at multiple levels of detail through two distinct processing streams, namely a Low-Level Feature Projection (LLFP) and a High-Level Feature Projection (HLFP) stream, for varying feature resolutions - enabling precise segmentation of anatomy and surgical instruments. We evaluated FASL-Seg on surgical segmentation benchmark datasets EndoVis18 and EndoVis17 on three use cases. The FASL-Seg model achieves a mean Intersection over Union (mIoU) of 72.71% on parts and anatomy segmentation in EndoVis18, improving on SOTA by 5%. It further achieves a mIoU of 85.61% and 72.78% in EndoVis18 and EndoVis17 tool type segmentation, respectively, outperforming SOTA overall performance, with comparable per-class SOTA results in both datasets and consistent performance in various classes for anatomy and instruments, demonstrating the effectiveness of distinct processing streams for varying feature resolutions.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06159
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FASL-Seg: Anatomy and Tool Segmentation of Surgical Scenes
Abdel-Ghani, Muraam
Ali, Mahmoud
Ali, Mohamed
Ahmed, Fatmaelzahraa
Arsalan, Muhammad
Al-Ali, Abdulaziz
Balakrishnan, Shidin
Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
I.4.6; I.4.8; J.3
The growing popularity of robotic minimally invasive surgeries has made deep learning-based surgical training a key area of research. A thorough understanding of the surgical scene components is crucial, which semantic segmentation models can help achieve. However, most existing work focuses on surgical tools and overlooks anatomical objects. Additionally, current state-of-the-art (SOTA) models struggle to balance capturing high-level contextual features and low-level edge features. We propose a Feature-Adaptive Spatial Localization model (FASL-Seg), designed to capture features at multiple levels of detail through two distinct processing streams, namely a Low-Level Feature Projection (LLFP) and a High-Level Feature Projection (HLFP) stream, for varying feature resolutions - enabling precise segmentation of anatomy and surgical instruments. We evaluated FASL-Seg on surgical segmentation benchmark datasets EndoVis18 and EndoVis17 on three use cases. The FASL-Seg model achieves a mean Intersection over Union (mIoU) of 72.71% on parts and anatomy segmentation in EndoVis18, improving on SOTA by 5%. It further achieves a mIoU of 85.61% and 72.78% in EndoVis18 and EndoVis17 tool type segmentation, respectively, outperforming SOTA overall performance, with comparable per-class SOTA results in both datasets and consistent performance in various classes for anatomy and instruments, demonstrating the effectiveness of distinct processing streams for varying feature resolutions.
title FASL-Seg: Anatomy and Tool Segmentation of Surgical Scenes
topic Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
I.4.6; I.4.8; J.3
url https://arxiv.org/abs/2509.06159