Video Dataset for Surgical Phase, Keypoint, and Instrument Recognition in Laparoscopic Surgery (PhaKIR)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rueckert, Tobias, Maerkl, Raphaela, Rauber, David, Klausmann, Leonard, Gutbrod, Max, Rueckert, Daniel, Feussner, Hubertus, Wilhelm, Dirk, Palm, Christoph
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912697592840192
author Rueckert, Tobias
Maerkl, Raphaela
Rauber, David
Klausmann, Leonard
Gutbrod, Max
Rueckert, Daniel
Feussner, Hubertus
Wilhelm, Dirk
Palm, Christoph
author_facet Rueckert, Tobias
Maerkl, Raphaela
Rauber, David
Klausmann, Leonard
Gutbrod, Max
Rueckert, Daniel
Feussner, Hubertus
Wilhelm, Dirk
Palm, Christoph
contents Robotic- and computer-assisted minimally invasive surgery (RAMIS) is increasingly relying on computer vision methods for reliable instrument recognition and surgical workflow understanding. Developing such systems often requires large, well-annotated datasets, but existing resources often address isolated tasks, neglect temporal dependencies, or lack multi-center variability. We present the Surgical Procedure Phase, Keypoint, and Instrument Recognition (PhaKIR) dataset, comprising eight complete laparoscopic cholecystectomy videos recorded at three medical centers. The dataset provides frame-level annotations for three interconnected tasks: surgical phase recognition (485,875 frames), instrument keypoint estimation (19,435 frames), and instrument instance segmentation (19,435 frames). PhaKIR is, to our knowledge, the first multi-institutional dataset to jointly provide phase labels, instrument pose information, and pixel-accurate instrument segmentations, while also enabling the exploitation of temporal context since full surgical procedure sequences are available. It served as the basis for the PhaKIR Challenge as part of the Endoscopic Vision (EndoVis) Challenge at MICCAI 2024 to benchmark methods in surgical scene understanding, thereby further validating the dataset's quality and relevance. The dataset is publicly available upon request via the Zenodo platform.
format Preprint
id arxiv_https___arxiv_org_abs_2511_06549
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video Dataset for Surgical Phase, Keypoint, and Instrument Recognition in Laparoscopic Surgery (PhaKIR)
Rueckert, Tobias
Maerkl, Raphaela
Rauber, David
Klausmann, Leonard
Gutbrod, Max
Rueckert, Daniel
Feussner, Hubertus
Wilhelm, Dirk
Palm, Christoph
Computer Vision and Pattern Recognition
Robotic- and computer-assisted minimally invasive surgery (RAMIS) is increasingly relying on computer vision methods for reliable instrument recognition and surgical workflow understanding. Developing such systems often requires large, well-annotated datasets, but existing resources often address isolated tasks, neglect temporal dependencies, or lack multi-center variability. We present the Surgical Procedure Phase, Keypoint, and Instrument Recognition (PhaKIR) dataset, comprising eight complete laparoscopic cholecystectomy videos recorded at three medical centers. The dataset provides frame-level annotations for three interconnected tasks: surgical phase recognition (485,875 frames), instrument keypoint estimation (19,435 frames), and instrument instance segmentation (19,435 frames). PhaKIR is, to our knowledge, the first multi-institutional dataset to jointly provide phase labels, instrument pose information, and pixel-accurate instrument segmentations, while also enabling the exploitation of temporal context since full surgical procedure sequences are available. It served as the basis for the PhaKIR Challenge as part of the Endoscopic Vision (EndoVis) Challenge at MICCAI 2024 to benchmark methods in surgical scene understanding, thereby further validating the dataset's quality and relevance. The dataset is publicly available upon request via the Zenodo platform.
title Video Dataset for Surgical Phase, Keypoint, and Instrument Recognition in Laparoscopic Surgery (PhaKIR)
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.06549