Surgical Triplet Recognition via Diffusion Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Daochang, Hu, Axel, Shah, Mubarak, Xu, Chang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929396706705408
author Liu, Daochang
Hu, Axel
Shah, Mubarak
Xu, Chang
author_facet Liu, Daochang
Hu, Axel
Shah, Mubarak
Xu, Chang
contents Surgical triplet recognition is an essential building block to enable next-generation context-aware operating rooms. The goal is to identify the combinations of instruments, verbs, and targets presented in surgical video frames. In this paper, we propose DiffTriplet, a new generative framework for surgical triplet recognition employing the diffusion model, which predicts surgical triplets via iterative denoising. To handle the challenge of triplet association, two unique designs are proposed in our diffusion framework, i.e., association learning and association guidance. During training, we optimize the model in the joint space of triplets and individual components to capture the dependencies among them. At inference, we integrate association constraints into each update of the iterative denoising process, which refines the triplet prediction using the information of individual components. Experiments on the CholecT45 and CholecT50 datasets show the superiority of the proposed method in achieving a new state-of-the-art performance for surgical triplet recognition. Our codes will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2406_13210
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Surgical Triplet Recognition via Diffusion Model
Liu, Daochang
Hu, Axel
Shah, Mubarak
Xu, Chang
Computer Vision and Pattern Recognition
Artificial Intelligence
Surgical triplet recognition is an essential building block to enable next-generation context-aware operating rooms. The goal is to identify the combinations of instruments, verbs, and targets presented in surgical video frames. In this paper, we propose DiffTriplet, a new generative framework for surgical triplet recognition employing the diffusion model, which predicts surgical triplets via iterative denoising. To handle the challenge of triplet association, two unique designs are proposed in our diffusion framework, i.e., association learning and association guidance. During training, we optimize the model in the joint space of triplets and individual components to capture the dependencies among them. At inference, we integrate association constraints into each update of the iterative denoising process, which refines the triplet prediction using the information of individual components. Experiments on the CholecT45 and CholecT50 datasets show the superiority of the proposed method in achieving a new state-of-the-art performance for surgical triplet recognition. Our codes will be released.
title Surgical Triplet Recognition via Diffusion Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2406.13210