Slot-BERT: Self-supervised Object Discovery in Surgical Video

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liao, Guiqiu, Jogan, Matjaz, Hussing, Marcel, Nakahashi, Kenta, Yasufuku, Kazuhiro, Madani, Amin, Eaton, Eric, Hashimoto, Daniel A.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911480194007040
author Liao, Guiqiu
Jogan, Matjaz
Hussing, Marcel
Nakahashi, Kenta
Yasufuku, Kazuhiro
Madani, Amin
Eaton, Eric
Hashimoto, Daniel A.
author_facet Liao, Guiqiu
Jogan, Matjaz
Hussing, Marcel
Nakahashi, Kenta
Yasufuku, Kazuhiro
Madani, Amin
Eaton, Eric
Hashimoto, Daniel A.
contents Object-centric slot attention is a powerful framework for unsupervised learning of structured and explainable representations that can support reasoning about objects and actions, including in surgical videos. While conventional object-centric methods for videos leverage recurrent processing to achieve efficiency, they often struggle with maintaining long-range temporal coherence required for long videos in surgical applications. On the other hand, fully parallel processing of entire videos enhances temporal consistency but introduces significant computational overhead, making it impractical for implementation on hardware in medical facilities. We present Slot-BERT, a bidirectional long-range model that learns object-centric representations in a latent space while ensuring robust temporal coherence. Slot-BERT scales object discovery seamlessly to long videos of unconstrained lengths. A novel slot contrastive loss further reduces redundancy and improves the representation disentanglement by enhancing slot orthogonality. We evaluate Slot-BERT on real-world surgical video datasets from abdominal, cholecystectomy, and thoracic procedures. Our method surpasses state-of-the-art object-centric approaches under unsupervised training achieving superior performance across diverse domains. We also demonstrate efficient zero-shot domain adaptation to data from diverse surgical specialties and databases.
format Preprint
id arxiv_https___arxiv_org_abs_2501_12477
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Slot-BERT: Self-supervised Object Discovery in Surgical Video
Liao, Guiqiu
Jogan, Matjaz
Hussing, Marcel
Nakahashi, Kenta
Yasufuku, Kazuhiro
Madani, Amin
Eaton, Eric
Hashimoto, Daniel A.
Image and Video Processing
Computer Vision and Pattern Recognition
Object-centric slot attention is a powerful framework for unsupervised learning of structured and explainable representations that can support reasoning about objects and actions, including in surgical videos. While conventional object-centric methods for videos leverage recurrent processing to achieve efficiency, they often struggle with maintaining long-range temporal coherence required for long videos in surgical applications. On the other hand, fully parallel processing of entire videos enhances temporal consistency but introduces significant computational overhead, making it impractical for implementation on hardware in medical facilities. We present Slot-BERT, a bidirectional long-range model that learns object-centric representations in a latent space while ensuring robust temporal coherence. Slot-BERT scales object discovery seamlessly to long videos of unconstrained lengths. A novel slot contrastive loss further reduces redundancy and improves the representation disentanglement by enhancing slot orthogonality. We evaluate Slot-BERT on real-world surgical video datasets from abdominal, cholecystectomy, and thoracic procedures. Our method surpasses state-of-the-art object-centric approaches under unsupervised training achieving superior performance across diverse domains. We also demonstrate efficient zero-shot domain adaptation to data from diverse surgical specialties and databases.
title Slot-BERT: Self-supervised Object Discovery in Surgical Video
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.12477